Hallucination Detection
Hallucination detection identifies generated claims that are unsupported, contradicted or otherwise unreliable under a specified evidence standard. It can use source comparison, consistency checks or trained evaluators, but the detector's scope must be explicit because contextual support, real-world factuality and repeated model agreement are different properties.
What it is
The term hallucination is used for several failures, including fabricated facts and claims unsupported by supplied context. A detector first needs an operational definition. Evidence-based approaches compare claims with source passages; consistency approaches compare multiple generations; specialized classifiers or model judges estimate support or contradiction. SelfCheckGPT is a particular sampling-based method motivated by inconsistent generations. These mechanisms do not all test the same thing. A claim repeated consistently can still be false, while a correct claim absent from the allowed context may fail a contextual-support criterion. Evaluation labels must reflect that distinction.
What the work involves
The practitioner decomposes outputs into checkable claims and specifies allowed evidence and the handling of uncertainty. It tests the detector on supported statements, contradictions, missing evidence and subtle qualifier changes. Precision and recall or equivalent error analysis matter because false alerts and missed errors have different consequences. Useful artifacts include claim annotations, detector configuration and a policy for review or correction. Detection should connect to an action, such as removing a claim or requesting more evidence, without assuming that a low-risk score verifies the entire response.
Illustrative example
An assistant describes a component's operating limits. The detector checks each limit against the supplied manual and finds that one temperature value has no supporting passage. The application retrieves further evidence or marks it unknown. In another case, repeated model samples agree on an obsolete value, illustrating why consistency alone is insufficient. Evaluation includes both cases so the team understands which errors the selected detector can catch and which require source version checks or expert review.
Limits and common mistakes
Detectors can fail when evidence is incomplete or a contradiction is subtle. Model-based checks may be vulnerable to misleading text and share the generator's blind spots. Sampling adds cost without producing independent facts. Claims of detection accuracy should name the hallucination definition, evidence setting and test distribution. No detector eliminates the need for source quality, calibrated uncertainty and application controls when a generated claim has significant consequences.
Prerequisites
RAG is the primary technique for mitigating hallucinations via grounding — understanding RAG informs detection strategies
Detecting hallucinations requires automated evaluation (faithfulness metrics, fact-checking) — eval frameworks are the detection tools
Related skills
- → is part of: AI Output Verification
- → is part of: LLM Evaluation Frameworks
Sources and further reading
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
Primary sampling-based detection method and its specific evaluation assumptions.
- Ragas faithfulness metric
Example of detecting unsupported claims through retrieved-context support.
Last updated: 2026-10-10