RAG evaluation measures how well a retrieval-augmented system finds appropriate evidence, uses it accurately, and satisfies the user's task. Evaluate those stages separately so a failure leads to a specific repair.
Numerical examples below are illustrative unless explicitly sourced.
Follow a wrong answer back through the RAG pipeline
A user asks whether a refurbished laptop can be returned after 20 days. The assistant says yes, citing a 30-day policy. The answer might fail for several different reasons: the refurbished policy was never ingested; parsing lost an exception; search found only the general policy; the reranker removed the relevant passage; context packing cut it off; or the model ignored it. “The RAG score is low” does not identify which repair is needed.
Start with the source record. Is there an authoritative policy for the product, region, and date? Check whether the user is allowed to see it. Then inspect the parsed text and searchable chunks. Only after confirming that the needed evidence exists in searchable form does it make sense to compare embedding models or ranking settings.
Hold one stage fixed while testing another. Give the generator a human-selected evidence packet to find out whether it can answer with sufficient context. Evaluate the retriever using labeled evidence independently of the generated prose. Compare reranked and original candidate lists to see whether reranking improved ordering or discarded useful evidence. This is the practical meaning of component-level evaluation.
Use frameworks without mistaking them for ground truth
Frameworks such as Ragas can automate parts of grading, including claim support, but the metric implementation, prompt, judge, and labels determine what the number means. Some metrics need references; others judge relationships between the question, answer, and context. Reference-free does not mean error-free, and a judge cannot establish facts that are absent from its evidence.
Use curated cases with source versions and review their labels. Synthetic questions can broaden coverage, but may overrepresent easy questions that echo document wording. Production feedback supplies useful failure cases, but positive feedback and silence are not verified correctness. Combine those sources while keeping their sampling purposes clear.
To control evaluation cost, run cheap deterministic checks first, cache only identical evaluation inputs under the same grader version, and use calibrated sampling or staged suites. Do not assign a perfect score because an answer contains “I don't know”: the refusal may be unnecessary or fail to answer a supported question.
Separate the evaluation layers
| Layer | What to verify | Example failure |
|---|---|---|
| Source/ingestion | Authoritative, current evidence exists and was preserved | A new policy was never indexed |
| Retrieval | Needed permitted evidence enters the selected set | General rule retrieved, exception missed |
| Context assembly | Evidence reaches the model with its scope intact | Packing drops the relevant qualifier |
| Generation | Claims follow from applicable evidence | Correct numbers are misinterpreted |
| User outcome | The response resolves the actual request | Accurate policy for the wrong country |
Read diagram source
flowchart LR
S[Known source snapshot] --> R[Candidate retrieval]
R --> P[Rerank and pack]
P --> G[Generate answer]
G --> O[Evaluate task outcome]
E[Reviewed evidence packet] --> G
L[Labeled relevant evidence] --> R
The reviewed packet tests generation independently of retrieval. Candidate labels test retrieval independently of the prose. Use both to diagnose a failure before changing the model.
Metrics with small examples
Suppose a labeled query has four relevant documents. The first five results contain three of them. Precision@5 = 3/5 = 60%. Recall@5 = 3/4 = 75%. Define relevance and the retrieval unit—document, passage, or evidence item—before counting. Incomplete labels make recall uncertain; an unjudged result is not automatically irrelevant.
Reciprocal rank is 1 / rank of the first relevant result: a first relevant result at rank three scores 1/3. Use zero when no relevant result is retrieved under the stated cutoff. Mean reciprocal rank (MRR) averages that value across queries. Normalized discounted cumulative gain (nDCG) accounts for graded relevance and position. Neither alone establishes that all evidence needed for a multi-part answer was retrieved. Check evidence coverage explicitly for those tasks.
Faithfulness asks whether answer claims are supported by retrieved context. Correctness asks whether the answer is correct relative to reliable reference evidence. Citation precision asks whether cited material supports the attached claims; citation coverage asks whether important claims have support. Ragas documents faithfulness as a support measure, not a universal truth test. Ragas faithfulness.
Calculate a graded ranking score
One common definition is DCG@k = Σ (2^relevance_i − 1) / log2(i + 1) for ranks starting at one. Divide by the best possible DCG for the same query and cutoff to obtain nDCG. State the gain and discount convention because variants exist. Information Retrieval textbook.
For three documents graded [2, 0, 1], DCG@3 is 3/1 + 0/log2(3) + 1/2 = 3.5. The ideal order [2, 1, 0] has DCG 3 + 1/log2(3) ≈ 3.6309, so nDCG@3 is approximately 0.964. This high score still does not prove that every required answer fact was covered. Define how zero-relevance queries are reported; do not divide by zero or silently count them as perfect retrieval.
Build the dataset from failure modes
Include common questions, rare but consequential questions, abbreviations, exact identifiers, languages, table/OCR problems, multi-document reasoning, conflicting versions, absent answers, and permission changes. Record the authoritative answer/evidence and the user's scope at the time of the test.
Synthetic questions help cover gaps but can be unnaturally easy or leak wording from the source. Validate them with people who understand the domain. Keep a protected holdout separate from examples used to tune chunking, prompts, and ranking.
A debugging ladder
- Search the raw source: is the answer actually there and current?
- Inspect parsing and chunks: was a table header or critical qualifier lost?
- Inspect candidates before reranking: if missing here, fix retrieval or filters.
- Inspect reranked candidates and packed context: did the needed passage get dropped?
- Inspect the generated claims and citations: did generation misread the evidence?
- Inspect the user's outcome: was the question ambiguous or the policy inapplicable?
Use ablations—change one stage and rerun matched cases—to learn which added complexity helps. A larger top-k can raise recall while adding distracting text, cost, and latency. A reranker cannot recover candidates never retrieved.
Release and live checks
Compare the candidate against baseline on the same queries. Set quality tolerances and severe-risk gates before evaluation; there is no universal faithfulness threshold that makes a system production-ready. Calibrate model graders against expert-labeled claims and inspect disagreement by slice.
Test ACL revocation, document deletion, stale caches, malicious passages, and safe abstention. Measure index freshness and complete-request latency, not only vector-search duration. Track representative live samples and user corrections; a thumbs-up rate is not unbiased correctness measurement.
Manager decisions
Name owners for source quality, search relevance, answer quality, and permissions. Budget annotation time and a recurring error-review meeting. When a metric improves, ask which real failure disappeared and whether any important slice became worse.
Grade four claims against explicit evidence
Assume the supplied current policy says: “Unopened standard products may be returned within 30 days. Refurbished products have a 14-day limit. Refunds use the original payment method.” The user asks about a refurbished laptop purchased 20 days ago.
| Generated claim | Evidence judgment | Explanation |
|---|---|---|
| Standard unopened products have a 30-day limit | Supported | Correctly describes the general rule |
| Refurbished products have a 14-day limit | Supported | Correctly preserves the exception |
| This refurbished laptop is eligible after 20 days | Contradicted | Applies the wrong conclusion despite stating the exception |
| The refund will arrive tomorrow | Unsupported | No processing-time evidence was supplied |
With equal claim weighting, faithfulness is supported / total = 2/4 = 0.50. Keep “contradicted” distinct from “not evidenced” in the failure record even if both count as unsupported for this score. The wrong eligibility conclusion is more consequential than a minor prose issue; the release rubric can make it a hard failure. Changing claim segmentation changes the denominator, so document the segmentation rule and calibrate it with experts.
Average precision is not precision at k
Suppose there are three relevant documents in the judged corpus, and the ranking is [relevant, irrelevant, relevant, irrelevant, relevant]. Precision@5 is 3/5 = 0.60. Precision at the relevant ranks is 1/1, 2/3, and 3/5; average precision is (1 + 2/3 + 3/5) / 3 = 0.7556. AP rewards putting relevant documents early; P@5 counts how many are in the first five. This query's reciprocal rank is 1 because the first result is relevant; it ignores the later evidence coverage.
For AP@k, explicitly state whether the denominator is all known relevant documents or min(total relevant, k); libraries differ. If labels are incomplete, report that limitation rather than treating all unjudged documents as proven irrelevant. Mean AP averages query-level AP; it is not the AP of one pooled list.
One grader contract
Use a versioned contract like this illustrative result, with quoted evidence spans or stable span IDs verified against the input:
{
"case_id": "refurbished-20-days",
"rubric_version": "claim-support-v3",
"source_revision": "returns-19",
"claims": [
{"id": "c1", "verdict": "supported", "evidence_ids": ["p1"]},
{"id": "c2", "verdict": "supported", "evidence_ids": ["p2"]},
{"id": "c3", "verdict": "contradicted", "evidence_ids": ["p2"]},
{"id": "c4", "verdict": "insufficient_evidence", "evidence_ids": []}
],
"faithfulness": 0.5,
"critical_error": "wrong_eligibility"
}
The judge receives the user question, source packet, answer, and rubric; it must not execute instructions embedded in any of them. Validate the JSON and evidence IDs, recompute the arithmetic outside the model, and allow an unjudgeable result. Calibrate against expert labels by claim type, language, and severity. For pairwise judging, blind model identities and vary answer order; for absolute scoring, inspect verbosity and style bias. Version the judge prompt/model, dataset, and reference evidence together.
Work a calibrated evaluation budget
Assume 100,000 answers/day. A frontier judge at a hypothetical $0.005/answer costs $500/day. Instead, run cheap deterministic contracts on all answers, sample 10,000 at known probability for a calibrated small judge at $0.0005 ($5), and escalate 1,000 uncertain or severe cases at $0.005 ($5). Add an independent random 1,000-answer frontier audit ($5) to detect blind spots in the routing. Model grading is then $15/day under these assumptions; deterministic infrastructure and human review are additional costs.
The targeted escalation set cannot estimate population error rates by itself. Keep the random sample's results separate or use justified sampling weights. Include fixed high-risk regression suites even if those cases are rare in traffic. If expert calibration shows the small judge misses consequential errors, change the routing or keep the stronger judge for that slice. The saving is acceptable only if the measurement still detects the regressions that matter.
Understand how response relevancy is estimated
One Ragas response-relevancy method generates plausible questions from the answer, embeds those questions and the user's actual question, and averages their cosine similarities. If three similarities are 0.9, 0.8, and 0.7, the mean is 0.8. This estimates whether the answer could address the intended question; it does not verify factual truth. Pin the metric implementation, judge, embedding model, and settings. See the Ragas calculation.
For “What is the return deadline and is there a fee?”, an answer describing only the deadline is incomplete even if the generated questions are semantically close to the user's request. Grade the required fee information explicitly. Conversely, a fluent answer giving a false deadline can be highly relevant. Keep relevance, completeness, claim support, and reference correctness as distinct checks. Cosine similarity itself can be negative, so do not assume every implementation's raw similarity is mathematically confined to 0–1.
For each test case, save the question, supplied context, answer, reference where required, individual verdicts, and scoring versions. Run the same cases for candidate and baseline; inspect disagreements and failures before averaging. Validate structured grader outputs and treat missing judgments separately. This record makes a lower score diagnosable: the retriever missed evidence, the answer omitted a required part, or the grader misread the response.
Interview questions with developed answers
Q1: Users report wrong RAG answers. How do you diagnose and fix them?
Sample answer: I collect examples with the query, identity scope, source version, retrieved passages, packed context, and answer. I first verify that the authoritative evidence existed and was correctly parsed. Then I test whether retrieval found it, whether reranking and packing preserved it, and whether generation used it correctly. I fix the observed stage rather than immediately swapping models. I add the failure to a regression suite and evaluate nearby cases so the repair does not break another slice. Permission and freshness failures receive separate attention because a plausible answer can still expose forbidden or obsolete information.
Follow-up: What if perfect supplied context fixes the answer? That points upstream toward retrieval or evidence preparation, though live integration still needs testing.
Q2: How do you evaluate RAG without ground-truth answers?
Sample answer: I begin with properties we can inspect: whether citations exist and support the claims, whether the answer addresses the question, and whether access and freshness rules hold. I use calibrated human or model review for semantic judgments and explicitly state the limitations. In parallel, I build a small expert-reviewed set of representative questions and required evidence. Synthetic cases help expand coverage but need validation. I would not describe reference-free faithfulness as overall correctness, because an answer can be supported by a bad source.
Follow-up: What should experts label first? High-value decisions and common failure categories where better evidence will change the system design.
Q3: Your RAG evaluation pipeline costs $500 per day. How do you reduce it?
Sample answer: I attribute the bill to suites, metrics, model calls, token lengths, and repeated unchanged cases. I remove redundant grading, cache exact inputs with the grader and rubric version, and use smaller calibrated graders only where agreement is adequate. I separate fast change-specific checks from broader scheduled suites while retaining required risk coverage. For production estimates, I sample with known probabilities and keep targeted incident investigations separate. I measure whether the cheaper evaluation still detects meaningful regressions; reducing the bill by making the measurement blind is not a useful optimization.
Follow-up: Can you reuse a score after the context changes? No; changed evidence changes the evaluation input.
Q4: High retrieval recall but poor answers—what does that tell you?
Sample answer: Recall says labeled relevant evidence was returned under a particular definition. It does not prove the evidence reached the model intact, was current, or covered every required fact. I inspect context packing, duplicates, contradictions, ordering, and citation behavior, then test the generator with a controlled evidence packet. I also verify the recall labels and unit. A passage can be relevant without being sufficient. Increasing top-k again may add noise and cost instead of addressing the downstream failure.
Follow-up: When would you inspect precision? When irrelevant context may be crowding out or distracting from the needed evidence.
Q5: What would make an evaluation result suitable for a release gate?
Sample answer: The cases, labels, source snapshot, configuration, and grader must be versioned and relevant to the intended users. The gate should distinguish critical policy violations from quality preferences and report slices with enough evidence for the decision. I would compare the candidate to a meaningful baseline, inspect regressions, and predefine acceptable tolerances. Passing offline checks allows the next controlled exposure stage; it does not replace production monitoring. The release owner must understand both what was tested and what uncertainty remains.
Follow-up: Are generic faithfulness thresholds portable? No; their meaning depends on the task, metric, judge, and consequences.
60-second interview answer
I evaluate retrieval, generation, and the final user outcome separately. First, can the retriever find relevant, current evidence the user is allowed to see? Second, does the answer accurately use that evidence and cite the right source? Third, does it solve the question, including abstaining when evidence is missing? I use labeled cases with answerable, unanswerable, conflicting, stale, and permission-restricted examples. I inspect failures stage by stage, calibrate semantic graders, and keep latency, cost, freshness, and access-control checks in the release gates.
Final notes
| Memorize the distinction | Do not infer |
|---|---|
| Recall measures labeled evidence found | The answer used all required evidence |
| Faithfulness measures support from supplied context | The source is current or true |
| Relevance measures fit to the question | All requested facts were answered correctly |
| Citation support connects claims to sources | A real URL supports every adjacent statement |
| A calibrated judge provides a useful measurement | The judge cannot be wrong |
Interview tip: Name the failing stage, the evidence needed to test it, and the release decision the metric supports. Keep critical-error gates and sampled population estimates separate.