Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

LLM evaluation: measure the behavior required by the product

By Anup Rai14 min readReviewed September 2026

Evaluation is the systematic assessment of a system against specified criteria. An LLM evaluation runs defined tasks, records outputs and observable effects, grades them using a stated method, and summarizes the evidence for a decision. The unit being evaluated may be a model, a retrieval component or the complete application.

A benchmark is a standardized evaluation used for comparison. A metric is a defined measurement. A rubric specifies the criteria and scoring rules for judgments. A grader applies those rules. A high metric value is useful only if it measures a property the product needs.

Conventional machine learning also deals with uncertain labels, subjective objectives and changing populations. LLMs add open-ended responses and tool behavior; they do not make ordinary experimental discipline obsolete.

Interview scope: release a policy assistant

The assistant answers questions from permitted policy documents. It may explain an exception-review process but cannot approve a refund. Numerical examples below are illustrative.

Functional requirements

  1. Run a versioned case set against the baseline and candidate.
  2. Capture answers, permitted evidence, tool outcomes and the configuration used.
  3. Apply exact checks, semantic rubrics and expert review where each is appropriate.
  4. Report improvements, regressions, severe failures and uncertainty by task slice.
  5. Support a release decision and turn diagnosed production failures into regression cases.

Non-functional requirements

  1. Keep evaluation data private and isolated from production effects.
  2. Prevent development/test leakage and preserve reproducible configurations.
  3. Bound execution, judging, review time and spending.
  4. Treat missing results and grader errors explicitly rather than silently dropping them.
  5. Preserve task diversity, including languages, difficult cases and rare consequential failures.

Start with representative cases, a small explicit rubric and a simple baseline. Add learned judges and large suites after understanding the failures they need to detect.

Separate dimensions before calculating a score

Dimension Question Counterexample
Correctness Does the answer match the applicable facts or specification? Fluent explanation uses the wrong return window
Faithfulness/grounding Are material claims supported by the supplied evidence? Correct fact has no support in the provided policy
Relevance Does it address the user's request? Accurate history of the company instead of return guidance
Completeness Does it cover required parts and qualifications? Gives the deadline but omits the applicable exception
Coherence and clarity Is the explanation understandable and internally consistent? Correct sentences contradict one another
Conciseness Is detail appropriate to the task? Short answer omits a necessary condition
Safety and authorization Are prohibited content and effects avoided? Correct refund amount sent to an unauthorized recipient
Helpfulness/task outcome Can the user achieve the intended goal? Safe refusal for every eligible question

For code, test execution, specified behavior, security and maintainability separately. For summaries, check key-fact coverage, contradictions and useful compression. A source-supported answer may still be false if the source is wrong, stale or inapplicable.

Build the evaluation contract

  1. Define the decision: diagnose a component, compare systems or approve a release.
  2. Define each case: input, permissions, source/environment state, expected behavior, forbidden effects and slice labels.
  3. Split by purpose: development cases, protected release holdout and targeted risk suite. Keep related conversations and near-duplicates together when splitting.
  4. Version the run: dataset, model revision/settings, prompts, tools, index snapshot, policy and grader.
  5. Record every outcome: success, system failure, timeout, grader failure or unavailable evidence.
  6. Analyze paired changes: run the same cases against both systems; inspect newly broken cases as well as newly fixed ones.

A risk suite can intentionally overrepresent dangerous requests. It cannot be naively mixed into a traffic-representative set and called population accuracy. Synthetic cases expand coverage but may inherit their generator's blind spots; audit them against real task requirements.

Repeated tuning on a holdout turns it into development data. A failure discovered in production is valuable regression coverage, but once engineers know it, it is no longer a fresh unseen test.

Choose the least ambiguous valid grader

Method Good use Limitation to state
Exact match Category, identifier or canonical result Normalize only allowed differences; case can matter
Keyword/presence checks Required field or literal clause A keyword can appear in a negation or an irrelevant sentence
Schema/parser Output contract Valid structure does not establish truthful content
Recalculation Arithmetic against authoritative inputs Wrong inputs still yield a wrong decision
ROUGE Reference-summary overlap ROUGE-N uses n-grams; ROUGE-L uses longest common subsequences; neither proves truth
Embedding similarity Semantic proximity or paraphrase signal Opposing decisions can have similar embeddings
Isolated executable tests Code behavior or external-state postconditions Passing tests proves only the tested properties
Calibrated model judge Semantic criteria at scale Bias, grader injection, cost and shared model errors
Expert review Consequential ambiguity and rubric design Time, disagreement and reviewer expertise

Do not run generated code with exec inside an evaluator holding production credentials. Use an isolated environment with resource limits and a protected verifier. The agent must not be able to edit its own test result or read hidden answers. Agent evaluation guidance.

A malformed judge result is a grader error, not automatically a bad answer or a pass. Report its rate and retry only within a budget. For 200 scheduled cases, if 180 pass, ten fail and ten have missing judgments, report 90% confirmed passes over all cases and 5% unknown, alongside 180/190 among graded cases. Do not quietly report 94.7% as if every case was assessed.

Diagnose RAG at the stage that failed

Measure Definition used here What it misses
Precision@k Relevant retrieved items divided by k, for a full top-k list Relevant documents outside the returned list
Recall@k Relevant items retrieved in top k divided by all labeled relevant items Depends on completeness of relevance labels
Reciprocal rank 1 divided by rank of the first relevant result; zero if none Other needed passages and their order
MRR Mean reciprocal rank across queries Coverage of multiple evidence requirements
Claim support Supported factual claims divided by assessed factual claims Whether the source itself is correct
Citation validity Citation identifies a real permitted source/span Whether that span supports the attached claim
Citation support/completeness Cited evidence supports claims and required claims are cited Overall usefulness and correctness need separate checks

If ranks 2 and 4 of five returned chunks are relevant and there are four labeled relevant chunks overall: precision@5 = 2/5 = 0.4, recall@5 = 2/4 = 0.5, reciprocal rank = 1/2 = 0.5. Deduplicate source units and define what a relevant unit means. Short lists, no relevant items, no factual claims and unanswerable questions need explicit scoring conventions.

Ragas offers multiple metric implementations; its rank-sensitive Context Precision variants are not simply interchangeable with the top-k fraction above. Its current faithfulness guide recommends the collections API (ragas.metrics.collections) over the legacy API. Pin the library and evaluator, and record which inputs each metric needs. A reference-free implementation still depends on its judge's validity. Ragas metric catalog, current faithfulness implementation guide.

Continue with RAG evaluation for retrieval failures, evidence packing and answer-level diagnosis.

Calibrate the judge

  1. Write observable criteria and examples of pass, fail and borderline behavior.
  2. Obtain independent labels from qualified reviewers and investigate disagreement.
  3. Compare the judge with those labels by error type and important slice.
  4. Randomize pairwise answer position and hide candidate identity where feasible.
  5. Test verbosity, formatting, family preference and instructions embedded in graded content.
  6. Version and recheck the judge after model, prompt, policy or traffic changes.

A pairwise judge should support a tie or insufficient evidence when the rubric permits it. Parse an explicit candidate ID; do not choose A merely because the first ten characters contain the letter A. Swapping positions and obtaining the same choice is a useful consistency check, not proof of calibrated confidence.

For binary acceptability, report missed unacceptable answers and false rejection of acceptable ones. For ordinal scales, consider weighted agreement measures with a justified weighting rule. Raw agreement and Cohen's kappa measure reviewer agreement, not correctness; arbitrary universal labels such as “0.8 means safe” hide domain and prevalence effects.

Read the numbers honestly

Illustration: 180 successes in 200 independent representative cases estimates 90% success. A Wilson 95% interval is about 85.1%–93.4%. Clustered conversations or repeated trials on the same tasks need analysis that respects those dependencies. Sample size depends on the uncertainty and risk you need to resolve, not a universal minimum of 100 or 200 cases.

With zero failures in n independent Bernoulli trials, the one-sided 95% upper bound is 1 − 0.05^(1/n), approximately 3/n for sufficiently large n. At n = 300 it is about 0.99%. Zero observed severe failures does not establish zero risk.

Suppose a candidate fixes 12 cases and breaks eight. The net gain of four hides severity and uncertainty. Analyze paired differences, repeated-run variability and important groups. A three-point gap is not automatically noise; a one-point gap is not automatically real. Set the comparison, allowable regressions and stopping rule before inspecting results. Repeated unplanned significance tests and selective reporting distort conclusions.

Design the pipeline and control its cost

Architecture / visual model
flowchart LR C[Versioned cases and reset environment] --> B[Run baseline and candidate] B --> R[Store outcomes and observable effects] R --> D[Exact and executable checks] D --> J[Semantic grading and expert audit] J --> A[Paired results, slices, uncertainty and cost] A --> G{Release criteria met} G -->|No| F[Diagnose and improve] G -->|Yes| P[Bounded rollout and live outcomes] P --> F
Read diagram source
flowchart LR
    C[Versioned cases and reset environment] --> B[Run baseline and candidate]
    B --> R[Store outcomes and observable effects]
    R --> D[Exact and executable checks]
    D --> J[Semantic grading and expert audit]
    J --> A[Paired results, slices, uncertainty and cost]
    A --> G{Release criteria met}
    G -->|No| F[Diagnose and improve]
    G -->|Yes| P[Bounded rollout and live outcomes]
    P --> F

Use durable queued work for background grading when results must survive process exits. Deduplicate by run/case/candidate/trial/grader version, store partial status, and retry failed work without overwriting valid evidence. Test the grader and environment as well as the candidate.

Illustrative budget: 1,000 cases × two candidates × three trials × $0.006 generation cost = $36. One $0.002 judgment per output adds $12. Reviewing 100 outputs for two minutes at $45/hour adds $150, for $198 before infrastructure. Smaller learned judges may help on validated slices; there is no universal traffic volume above which frontier judging becomes unaffordable or universally obsolete.

Added control Benefit Cost/limit
Repeated trials Reveals unstable completion More runs; correlated trials still need care
More expert labels Better rubric and disagreement diagnosis Specialist time and ongoing calibration
Cheap judge plus escalation Lower routine grading cost Confidently wrong routing requires independent audit
Isolated shadow replay Tests current inputs without intended production effects May not reproduce all live conditions
Bounded canary Measures real user outcomes Requires containment, rollback and sufficient exposure

Release gates and production learning

Separate hard requirements from quality tradeoffs. A high helpfulness mean cannot offset known cross-tenant exposure. Name the release owner, exception rules and rollback conditions. Monitor task completion, complaints, abandonment, latency, cost and selected expert review; feedback buttons are selected feedback, not a representative truth label.

When scores change, check the population, source data and grader before blaming the model. An average shift of 0.1 is an effect-size threshold, not a statistical test. Use the same reference cases to diagnose grader drift. Refresh coverage while preserving traceability to prior releases.

One complete evaluation record and a scored rubric

A reproducible record binds the input, evidence, configuration, output, and judgment. This fictional example evaluates a policy answer, not the model's hidden reasoning.

{
  "case_id": "returns-refurbished-20d",
  "dataset_version": "support-test-12",
  "input": "Can I return a refurbished laptop bought 20 days ago?",
  "evidence": [
    {"id": "policy-19-p2", "text": "Refurbished returns: 14 days."},
    {"id": "policy-19-p3", "text": "Customers outside the standard window may request manual exception review; approval is not guaranteed."}
  ],
  "expected_behavior": "Explain ineligibility under the supplied rule; offer permitted escalation.",
  "release": "support-candidate-b",
  "output": "The policy allows 14 days, so this purchase is outside that window. I can explain the exception-review process.",
  "rubric_version": "support-4",
  "scores": {"correctness": 2, "evidence": 2, "next_step": 1},
  "critical_violation": false,
  "grader": "expert-review-7",
  "latency_ms": 1850,
  "total_cost_usd": 0.006
}

Correctness and evidence each use 0 (wrong/absent), 1 (partly satisfied without a material contradiction), or 2 (all required material points satisfied); next step is 0 or 1. The record scores 5/5 because the 14-day rule supports the eligibility decision and the supplied exception-review clause supports the offered next step. Without that second clause, the escalation offer would be unsupported and should not receive full evidence credit. “The limit is 30 days, so I issued a refund” fails correctness and evidence and may trigger a prohibited-action gate. Keep severity outside the weighted mean. Validate the source and applicable user state rather than accepting the reference label as infallible.

Calculate reviewer agreement

Two reviewers label 100 answers acceptable/unacceptable:

Reviewer B acceptable Reviewer B unacceptable Total
Reviewer A acceptable 60 10 70
Reviewer A unacceptable 5 25 30
Total 65 35 100

Observed agreement is (60 + 25)/100 = 0.85. Agreement expected from the marginals is 0.70 × 0.65 + 0.30 × 0.35 = 0.56. Cohen's kappa is (0.85 − 0.56)/(1 − 0.56) ≈ 0.659. This is an agreement statistic, not proof either reviewer is correct. Prevalence, label imbalance, sampling, and ambiguous categories affect it. Review disagreements, clarify the rubric, and relabel an independent sample; do not simply force consensus until the number looks good.

A memory evaluation needs separate operations

A user says “I used to live in Boston; I now live in Seattle.” Extraction should distinguish the current fact from historical context. Update should supersede the active Boston fact with Seattle while retaining appropriate provenance/history under policy. Read should retrieve the current permitted fact when relevant. If extraction was correct but the answer still says Boston, inspect update conflict resolution and read selection separately.

Now the user deletes the location preference. Test deletion through facts, summaries, caches, and paused workflows. A restored checkpoint must not reintroduce the deleted fact. Grade unnecessary recall too: remembering the correct city does not mean every answer should mention it.

Layer grading without equating different judges

Architecture / visual model
flowchart TD O[Outputs and observable task traces] --> D[Deterministic contracts on all cases] D --> S[Calibrated small judge on suitable slices] S --> F[Stronger judge for ambiguity or disagreement] F --> H[Expert review for consequential or unresolved cases] O --> A[Independent random audit sample] A --> H H --> C[Calibration and regression cases]
Read diagram source
flowchart TD
    O[Outputs and observable task traces] --> D[Deterministic contracts on all cases]
    D --> S[Calibrated small judge on suitable slices]
    S --> F[Stronger judge for ambiguity or disagreement]
    F --> H[Expert review for consequential or unresolved cases]
    O --> A[Independent random audit sample]
    A --> H
    H --> C[Calibration and regression cases]

A process reward model (PRM) scores intermediate steps under a learned process-supervision setup, often in reasoning tasks. An agent trajectory judge evaluates observable tool choices, order, permissions, and final state. They are not interchangeable merely because both inspect steps. Do not claim access to hidden internal reasoning; evaluate the evidence actually available. A correct final answer can hide a forbidden tool action, while a verbose rationale can be wrong.

The layer order is an operating choice, not a theorem that a larger judge is always more accurate. Keep an independent audit sample to detect cases the cheap stage confidently misroutes, measure agreement with experts by slice, and include human review in the cost budget.

Distinguish occasional success from reliable repetition

pass@k asks whether at least one of k attempts succeeds. pass^k asks whether all k attempts succeed under the benchmark's stated protocol. The first can describe a search process that gets several tries; the second stresses consistent completion. A user making one consequential request does not automatically get the benefit of an oracle selecting a successful attempt.

With independent attempts of identical success probability p = 0.8, the illustrative probabilities at k = 4 are pass@4 = 1 − (1 − 0.8)^4 = 99.84% and pass^4 = 0.8^4 = 40.96%. Real tasks differ in difficulty, attempts may be correlated, and benchmark estimators have their own sampling definitions. Measure repeated runs per task; do not raise an aggregate pass rate to a power and present it as an empirical result. The τ-bench paper uses repeated-trial consistency to assess agents.

Save task identity, environment reset, run count, seeds/settings where available, outcome, and prohibited effects. A low all-runs success rate tells you consistency is weak; it does not by itself prove whether the cause is poor recovery, ambiguous policy, a flaky tool, or unstable model decisions. Inspect trajectories to determine that cause. For costly or irreversible effects, run repetitions in isolated test environments.

Interview questions with developed answers

Q1: How would you evaluate a RAG system?

Sample answer: I separate evidence retrieval from answer generation. I first check whether the relevant, current, authorized source exists and whether the retriever returns the needed passage. Then I check whether the answer is supported by that passage, correct for the question, complete, and properly cited. I include questions with missing or conflicting evidence so clarification and abstention are tested. I compare to a baseline on held-out cases and report results by language, document type, and other important slices. Finally I measure user outcomes, latency, and full cost in a bounded rollout. An answer-level score alone does not reveal which component to fix.

Follow-up: Can a faithful answer still be wrong? Yes, if it accurately repeats a false, obsolete, or inapplicable source.

Q2: What are the limitations of LLM-as-judge?

Sample answer: A model judge is a fallible measurement component. It may prefer a longer answer, favor one position in a comparison, miss domain errors, or follow malicious instructions inside the candidate response. I define an explicit rubric, collect independent expert labels, measure disagreement by error type, and randomize presentation where relevant. I use deterministic checks for properties that can be verified directly. I version the judge and periodically recalibrate it. For high-impact ambiguous cases, I retain expert review rather than assuming a larger judge model removes the uncertainty.

Follow-up: Would using a different model family solve bias? It can reduce some shared preferences, but calibration is still required.

Q3: How do you evaluate when no single correct answer exists?

Sample answer: I define the properties an acceptable answer must satisfy. A summary can have many valid phrasings but must preserve key facts, avoid contradictions, and cover specified decisions. Domain experts can label those properties or compare responses under a shared rubric. I keep factual and safety checks separate from style preferences. I also measure whether users complete the intended task. The absence of one reference sentence changes the grading method; it does not mean we should replace evaluation with subjective impressions from a demo.

Follow-up: What if reviewers disagree? Inspect whether the rubric is ambiguous, the task has multiple legitimate goals, or specialized expertise is needed.

Q4: The score improved from 90% to 92%. Would you ship?

Sample answer: I would ask how many cases were tested, whether they represent the workload, and which cases changed. I would examine paired improvements and regressions, uncertainty, repeated-run variability, and severe failures. I would also verify that the same grader and data versions were used. If the quality evidence is sufficient and hard requirements pass, I would use a bounded rollout with operating and business metrics. A two-point average improvement is useful evidence only after its measurement conditions and consequences are understood.

Follow-up: What does zero observed severe failures prove? Only that none occurred in that sample, not that the true risk is zero.

Q5: How do you stop a team from gaming its evaluation?

Sample answer: I would avoid making one aggregate score the sole success criterion. We need protected holdouts, refreshed real-world cases, severity-based analysis, and production outcomes alongside development metrics. Engineers should be rewarded for exposing failures and improving user results, including choosing not to launch. I would investigate a widening gap between offline scores and complaints, inspect leakage and grader shortcuts, and rotate evaluation responsibilities. The goal is to make the evaluation a decision aid whose weaknesses are discussed openly, not a target people can improve without improving the product.

Follow-up: Should every production failure enter the holdout? It can become a regression case, but known cases then belong to development-visible coverage rather than an untouched holdout.

60-second interview answer

I start with the product decision the evaluation must support. I define success and unacceptable failures, build representative cases plus targeted risk cases, and compare the candidate against a meaningful baseline. Deterministic checks handle exact requirements; calibrated human or model graders handle semantic judgments. I report quality by slice with uncertainty, alongside latency and total cost. An average score cannot compensate for a severe safety regression. Evaluation continues after launch through sampled review, user outcomes, and controlled experiments, with a named owner for the release decision.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← AI governance and compliance: turn obligations into operated controls
NEXT LESSONAI observability: explain what happened to a user task →

Explore the diagram