AML.T0062
Discover LLM Hallucinations
Description
Adversaries may prompt large language models and identify hallucinated entities. They may request software packages, commands, URLs, organization names, or e-mail addresses, and identify hallucinations with no connected real-world source. Discovered hallucinations provide the adversary with potential targets to Publish Hallucinated Entities. Different LLMs have been shown to produce the same hallucinations, so the hallucinations exploited by an adversary may affect users of other LLMs.
Honesty-tier rationale
Discovery mostly probes the structure of real deployments; only the concept can be shown against a mock endpoint.