RAG Evaluation
RAG evaluation measures retrieval quality and answer accuracy to check whether a RAG system is actually helpful.
Prerequisites
Why Evaluate RAG?
A RAG system can look like it's working — fluent answers, plausible-sounding sources — while actually retrieving the wrong content, ignoring what it retrieved, or hallucinating despite having relevant context available. Evaluation is how you find out which of these is actually happening, instead of judging quality by whether the final answer merely reads well.
Key Idea
RAG has two things that can fail independently: retrieval (did we find the right content?) and generation (did the model use it correctly?). Evaluating only the final answer conflates the two.
What to Measure
- Retrieval relevance
- Did the system retrieve chunks that actually contain information relevant to the query? Measured independently of what the model does with them.
- Groundedness (faithfulness)
- Does the generated answer actually rely on the retrieved content, or does it contain claims not supported by any retrieved chunk? This directly targets hallucination — an ungrounded answer may be fluent and even correct by coincidence, but isn't verifiably supported by what was retrieved.
- Answer relevance
- Does the final answer actually address the user's question, independent of whether it's grounded in the retrieved content?
- Citation / source attribution
- When an answer references specific facts, can it point to which retrieved chunk supports each one? Good attribution makes an answer auditable — a user or reviewer can check the claim against its source instead of trusting the model's word.
How Evaluation Is Done
- Golden test sets — a curated set of representative questions with known-good expected answers or expected source documents, run against the system to measure retrieval and answer quality directly.
- LLM-as-judge — using a separate model call to score groundedness or relevance against the retrieved context, useful for scaling evaluation beyond what humans can manually review.
- Human review — spot-checking a sample of real interactions, especially valuable for catching failure modes an automated judge might miss.
- Regression testing — re-running the same test set whenever the retrieval pipeline, chunking strategy, or prompt changes, to catch quality regressions before they reach users.
Important
A wrong answer is often caused by bad retrieval, not a weak model. Diagnosing which stage failed — retrieval or generation — is usually more useful than only measuring the final answer.
Common Mistakes
Evaluating only the final answer
This conflates retrieval failures with generation failures, making it hard to know what to actually fix.
Never checking groundedness
A fluent, plausible answer can still be unsupported by anything actually retrieved — this is a distinct failure mode from simple wrongness.
Relying entirely on LLM-as-judge without any human review
An automated judge can share blind spots with the model it's judging — periodic human review catches what automation misses.
Testing once and never again
Changes to chunking, retrieval, or prompts can silently regress quality — evaluation should be a repeatable process, not a one-time check.
Using an unrepresentative test set
A test set that doesn't reflect real user questions can hide the failures that actually matter in production.
Interview Question
How would you evaluate whether a RAG system is actually working well?
I'd separate retrieval quality from generation quality, because they fail independently — a wrong answer is often caused by bad retrieval, not a weak model. For retrieval, I'd measure whether the system finds chunks actually relevant to representative test questions. For generation, I'd measure groundedness — whether the answer's claims are actually supported by what was retrieved — and answer relevance, whether it addresses the question at all. In practice that means a golden test set, LLM-as-judge scoring for groundedness at scale, periodic human review to catch what automated judging misses, and re-running the whole evaluation whenever the pipeline changes, so regressions get caught before they reach users.
What an interviewer may ask next
- Why is it important to evaluate retrieval and generation separately rather than only the final answer?
- What is groundedness, and why does it specifically target hallucination?
- Why would you still want human review even if you have an LLM-as-judge pipeline?
- How would you diagnose whether a wrong answer was caused by bad retrieval or bad generation?
Explain It in 30 Seconds
RAG evaluation measures retrieval quality and generation quality separately, since they fail independently — a wrong answer is often bad retrieval, not a weak model. Key metrics are retrieval relevance, groundedness (whether the answer is actually supported by what was retrieved), and answer relevance. In practice, teams use a golden test set, LLM-as-judge scoring to scale beyond manual review, periodic human spot-checks, and repeat the whole evaluation whenever the pipeline changes to catch regressions.