Hallucination
Hallucination is when a model generates confident, fluent output that is factually incorrect or unsupported.
What Is Hallucination?
A language model generates text by predicting the most likely next token, based on patterns learned during training — it has no built-in mechanism to check whether a statement is actually true. Hallucination is what happens when that fluent, confident-sounding prediction is wrong: a fabricated citation, an invented API method, a plausible-sounding but incorrect fact.
Key Idea
A model does not know the difference between "confident and correct" and "confident and wrong" — both can look identical in its output.
This isn't a bug that only shows up occasionally in a broken model — it's a direct consequence of how generation works. The model is always producing its best statistical guess at the next token, whether or not that guess happens to be true.
Why It Happens
- Knowledge gaps — the model was never trained on the specific fact being asked about, but still generates a plausible-sounding answer instead of refusing.
- Outdated training data — the model confidently states something that was true at training time but has since changed.
- Ambiguous or underspecified prompts — the model fills in missing details with a plausible guess rather than asking for clarification.
- Pattern completion — the model is very good at continuing a plausible-sounding pattern, such as a citation format, even when it has to invent the specific details to fit it.
Reducing Hallucination
You cannot fully eliminate hallucination from a language model, but several practical techniques reduce how often it happens and how much damage it does:
- Retrieval-Augmented Generation (RAG) — ground the answer in retrieved, verifiable source material instead of relying purely on what the model memorized.
- Lower temperature — reduces the chance of the model wandering into a less likely, potentially fabricated continuation.
- Asking the model to cite sources or say "I don't know" — explicitly permitting uncertainty reduces pressure to always produce a confident-sounding answer.
- Human review for high-stakes output — treating model output as a draft that a person verifies, rather than a final answer.
Common Mistakes
Trusting confident-sounding output as automatically correct
Fluency and confidence are properties of the language, not evidence that the underlying claim is true.
Assuming RAG eliminates hallucination entirely
RAG reduces hallucination by grounding answers in retrieved content, but the model can still misread or misstate what was retrieved.
Assuming hallucination only affects obscure facts
Models can hallucinate confidently about well-known topics too, especially with ambiguous or leading prompts.
Shipping ungrounded LLM output directly into high-stakes decisions
Medical, legal, financial, or safety-critical use cases need verification or human review, not a raw model response.
Interview Question
What is hallucination in the context of LLMs, and how would you reduce it in an application?
Hallucination is when a model generates fluent, confident output that is factually wrong or unsupported, because generation is fundamentally next-token prediction with no built-in fact-checking — a wrong guess looks exactly as confident as a right one. It happens most with knowledge gaps, outdated training data, or ambiguous prompts that get filled in with a plausible guess. You can't eliminate it, but you can reduce it: ground answers in retrieved, verifiable content with RAG, lower the temperature, explicitly allow the model to say it doesn't know, and add human review for high-stakes output.
What an interviewer may ask next
- Why can't hallucination be fully eliminated from a language model?
- How does RAG reduce hallucination, and why doesn't it eliminate it completely?
- Why might a model hallucinate more on an ambiguous prompt than a clear one?
- What would you do differently for a high-stakes use case where a hallucination could cause real harm?
Explain It in 30 Seconds
Hallucination is when a model produces fluent, confident text that's factually wrong, because it's always generating its best statistical guess at the next token with no built-in way to check truth. It shows up most with knowledge gaps, outdated training data, or vague prompts. You reduce it — not eliminate it — by grounding answers in retrieved sources with RAG, lowering temperature, allowing the model to express uncertainty, and adding human review for anything high-stakes.