Learnastra AI SYSTEM DESIGNAnup Rai

Complete design interview

Design an Evaluation Gate for AI Releases

By Anup Rai18 min readReviewed September 2026

Hypothetical interview scenario. Workloads, tolerances, costs and service targets are assumptions for this design. They are not Learnastra operating results or universal standards.

Interview focus: Decide whether a specific AI change is ready to merge and release, using reproducible evidence and an explicit treatment of uncertainty.

1. Define the problem and scope

Continuous integration (CI) automatically builds and tests proposed software changes. Continuous delivery keeps validated changes ready for release; continuous deployment also releases them automatically after the required checks. An evaluation gate is a required decision based on tests of the AI system's behavior. It complements ordinary software tests.

Design the gate for a 28-engineer team with four product teams and 50 AI-related pull requests (PRs) a week. The product answers questions about customer contracts using retrieval and tools. A shorter-answer prompt can omit a renewal exception while making the average style score improve. The gate must detect that consequential omission, explain the evidence and identify the exact build it assessed.

Clarify with the interviewer: Which failures must prevent release? Who labels correct contract interpretations? Does this gate allow merge, production exposure, or both? Can evaluation access customer data? What is the acceptable delay when evidence is inconclusive?

Functional requirements

  1. Test candidate and baseline behavior on versioned inputs, source documents, tool fixtures and scoring criteria.
  2. Run deterministic contract checks, representative quality comparisons and targeted severe-failure tests as separate suites.
  3. Select additional coverage from the changed components and affected customer domains.
  4. Validate semantic judges against expert labels and retain their calibration versions.
  5. Report pass, block or hold for more evidence, with changed examples and per-criterion results.
  6. Bind the decision to a build and evidence manifest; invalidate it when a relevant dependency changes.
  7. Promote an approved artifact through controlled deployment and retain a compatible rollback release.
  8. Maintain private traces, reviewer decisions, budgets and an auditable exception process.

Non-functional requirements

  1. Integrity: no merge based on a stale, incomplete, cancelled or untrusted check. A known critical tenant-data disclosure blocks regardless of aggregate quality.
  2. Latency: aim for p95 evaluation completion within 60 minutes of admission, leaving review time within a 90-minute PR-to-merge target. Report queue and execution time separately.
  3. Cost: default PR model spend below $40; explicit approval above that. A $1,200 full-run ceiling is an admission limit, not the predicted price.
  4. Privacy: tenant-approved data, minimum runner privileges and restricted access to customer examples and independent holdouts.
  5. Reproducibility: preserve the tested configuration and evidence. Model nondeterminism still requires measured repeated trials.
  6. Operations: failures remain visible and pending until resolved; quarterly methodology review is this scenario's governance requirement, not a universal legal mandate.

Out of scope: proving that no future failure can occur, letting a model approve its own security exceptions, and executing irreversible customer actions in test environments.

2. Establish a useful baseline

Begin with code tests plus a small, reviewed regression suite for known problems: correct currency, permitted tenant, valid tool arguments, real citation IDs and the renewal exception. Store expected properties rather than an exact prose answer. Run the candidate in an isolated environment, compare against the current production release and require a reviewer to inspect failures.

Baseline weakness Why it matters Next improvement and cost
Only known examples Misses new tasks and languages Add a representative sample and fresh error analysis; ongoing labeling work
One overall judge score Style can mask omitted obligations Separate critical requirements and quality dimensions; more reporting
Saved old score versus newly graded candidate A judge change looks like a product change Grade comparable outputs with the same rubric/version; extra calls
PR passes before another PR merges Their combined behavior is untested Evaluate the prospective merge revision; extra queue work
Green offline report means immediate broad release Fixtures miss external-service and live-traffic behavior Integration tests and canary exposure; slower rollout

A small suite is useful development feedback. Calling it evidence of population-wide quality requires a justified sampling plan.

3. Size the evaluation workload

Use a 4,000-case broader quality suite and a 500-case default PR suite as initial planning sizes, not statistical guarantees. Mandatory severe-risk and previously broken behaviors are separately identified; repeated or deliberately difficult examples must not be counted as independent representative observations.

Quantity Calculation Implication
Default work 500 cases × 2 builds × generation and grading 2,000 model calls per PR for one criterion
Default week 50 × 2,000 100,000 calls before full runs/retries
Sequential model time 2,000 × assumed 3 seconds 6,000 seconds of call time
Ideal 20-call concurrency 6,000 / 20 5 minutes lower-bound service time; dependencies, rate limits and long tails add time
Weekly execution demand 50 × assumed 12 runner-minutes 10 runner-hours; averages do not size a release-day burst
Ten simultaneous PRs 10 × 20 active calls Up to 200 calls competing for provider limits

Admit jobs against both request and token quotas. At 4,000 input and 600 output tokens, 20 calls each taking three seconds would demand roughly 1.84M combined tokens/minute. Providers may meter input and output separately. More workers cannot bypass those limits. Reserve capacity for the required merge checks, cancel superseded runs and cap retries.

4. Detailed architecture and evidence contracts

Architecture / visual model
flowchart TB PR[PR and prospective merge revision] --> PLAN[Trusted planner<br/>change scope and required suites] PLAN --> ADMIT[Budget and rate-limit admission] DATA[Versioned cases<br/>source snapshots and tool fixtures] --> ADMIT ADMIT --> RUN[Isolated baseline and candidate workers] RUN --> EXACT[Exact contract and security assertions] RUN --> SEM[Semantic graders<br/>pinned rubric and model] HUMAN[Independent expert labels<br/>calibration and disagreements] --> SEM EXACT --> EVIDENCE[Append-only evidence manifests] SEM --> EVIDENCE EVIDENCE --> GATE[Trusted gate<br/>identity, coverage and uncertainty] GATE -->|Block or hold| REPORT[Scoped report to developer] GATE -->|Pass| CHECK[Required check on exact revision] CHECK --> RELEASE[Artifact promotion<br/>shadow and canary stages] RELEASE --> MON[Live outcomes and rollback controller] MON --> DATA
Read diagram source
flowchart TB
    PR[PR and prospective merge revision] --> PLAN[Trusted planner<br/>change scope and required suites]
    PLAN --> ADMIT[Budget and rate-limit admission]
    DATA[Versioned cases<br/>source snapshots and tool fixtures] --> ADMIT
    ADMIT --> RUN[Isolated baseline and candidate workers]
    RUN --> EXACT[Exact contract and security assertions]
    RUN --> SEM[Semantic graders<br/>pinned rubric and model]
    HUMAN[Independent expert labels<br/>calibration and disagreements] --> SEM
    EXACT --> EVIDENCE[Append-only evidence manifests]
    SEM --> EVIDENCE
    EVIDENCE --> GATE[Trusted gate<br/>identity, coverage and uncertainty]
    GATE -->|Block or hold| REPORT[Scoped report to developer]
    GATE -->|Pass| CHECK[Required check on exact revision]
    CHECK --> RELEASE[Artifact promotion<br/>shadow and canary stages]
    RELEASE --> MON[Live outcomes and rollback controller]
    MON --> DATA

The planner, grader and check reporter run from a trusted version. A PR that can replace its own gate with return pass defeats the design. Separate untrusted candidate execution from the authority to read restricted holdouts or write a successful check. Use short-lived, scoped credentials, isolated runners and redacted artifacts. Do not expose deployment secrets to untrusted PR code or treat log masking as complete data protection. GitHub Actions secure use.

Practical implementation choices

Component Starting choice Why and when to change
CI control GitHub Actions plus protected required checks Fits the assumed PR workflow; keep the gate portable to other CI systems
Exact assertions Pytest and domain-specific validators Transparent expected outcomes; do not encode semantic truth as fragile string equality
Case and result storage Versioned JSONL in approved object storage plus a relational run catalog Cheap immutable evidence and queryable status; catalog and blobs need consistent publication
Traces and comparisons Langfuse with access controls and retention Supports inspecting evaluation experiments; storing a score does not enforce the release policy
Human annotation Restricted review UI or Argilla Independent labeling and disagreement resolution; budget domain-expert time

Langfuse evaluation documentation distinguishes offline experiments from online evaluation. Implement the required-check policy explicitly instead of assuming an observability dashboard blocks deployment.

Record Required content Contract
Case ID/version, input, expected properties, source permissions, slice, sampling weight and cluster ID Source and label versions are immutable for a run
Build manifest Code/artifact digest, prompts, model/config, index snapshot, tools and policies Describes the executed behavior, not only a Git SHA
Evaluation run Run ID, exact baseline/candidate manifests, suite, judge, repetitions, budgets and status Deduplicate triggers by the full run specification
Observation Case/build/trial IDs, output, assertion results, judge details, duration and billable usage Missing output is recorded, not silently excluded
Gate decision Required checks, coverage, per-slice deltas/intervals, policy version and decision Only trusted evaluator evidence can produce a pass
Exception Owner, reason, impacted criteria, exposure limit, compensating controls and expiry Explicitly distinguished from an ordinary pass
Release Artifact manifest, approved evidence digest, rollout state and prior compatible release Promotion checks identity again

A possible internal API is POST /eval-runs with immutable manifest IDs and an idempotency key. GET /eval-runs/{id} exposes queued, running, completed, failed or cancelled; completed evaluation is not synonymous with a passing gate. POST /release-decisions accepts trusted evidence references, not a candidate-supplied score.

When using GitHub's merge queue, run required checks on the merge group revision through its merge_group event. It includes the current base and preceding queued changes, so it differs from the original PR revision. A green check from the old revision is insufficient. Apply the equivalent contract in another CI provider. GitHub merge queues.

5. Build the right test sets and rubric

Separate three kinds of evidence

Suite How cases are selected What its score can tell you
Representative quality Permitted traffic sampled by a documented design Estimated performance for that population, using correct weights
Targeted regressions and attacks Known failures, synthetic attacks and important edge cases Whether these specified risks still fail; not their population frequency
Restricted assessment Fresh independent cases, access controlled and not repeatedly tuned against A less contaminated assessment of the selected candidate

Stratify by contract type, language and workflow where needed. If a rare category is oversampled, weight it back for the overall population estimate and still show its separate result. One example per category gives coverage of names, not reliable measurement. With a true independent 1% failure rate, ten random cases miss every failure with probability 0.99^10 ≈ 90.4%.

Keep serious historical regressions required while they remain applicable. Archive obsolete behaviors with a reason rather than deleting inconvenient failures. Rotate fresh assessment cases based on drift and exposure; “replace 10% quarterly” does not itself prevent overfitting. Repeated pass/fail feedback also leaks information about a holdout, even if raw examples are hidden.

Turn observed errors into assertions

Open coding means labeling observed failures with specific descriptions. Axial coding groups and relates those labels into useful families. Two reviewers apply the definitions to fresh examples and resolve ambiguity before scaling annotation.

Observed failure Failure family Test and owner
Current contract retrieved, old notice period used Wrong source applicability Check date and contract edition; retrieval/domain team
Renewal exception omitted Incomplete answer Rubric requires exception when condition applies; domain reviewers
Citation ID does not exist Invented evidence Exact ID membership plus semantic support check; application team
Refund currency differs from order Invalid action Exact pre-execution currency assertion; payments integration team
Correct answer contains another tenant's clause Unauthorized disclosure Isolated tenant fixtures and access tests; security and application teams

A citation ID resolving is necessary but does not establish that the cited text supports the claim. Likewise, valid JSON can still contain a wrong amount. Define refusal errors in both directions: failing to refuse a prohibited request and refusing a permitted one.

6. Check the judge before trusting its scores

An LLM judge grades an output against a rubric. It is a noisy measurement component, not ground truth. Use code where the result is objectively executable; use domain experts where correctness requires interpretation, and test whether the judge agrees for each relevant criterion.

  1. Create written labels with examples and counterexamples. Resolve human disagreement and retain uncertainty.
  2. Keep judge-development, selection and final assessment data distinct. A 60/20/20 split is an option, not a requirement; all splits need enough relevant examples.
  3. Pin grader prompt, model, rubric and generation settings. Randomize answer order in pairwise evaluation and test for position, verbosity and self-preference bias.
  4. Track false passes and false failures by domain. Overall accuracy and Cohen's kappa do not establish adequate recall for rare harmful errors.
  5. Recalibrate when the application, judge or population changes. Monthly fresh labeling is a cadence hypothesis; a material change requires immediate review.

Position and verbosity biases are documented in the original LLM-as-a-judge study. Two judges agreeing is useful diagnostic evidence, but they may share the same mistake. A larger model can help adjudicate cases without replacing expert labels.

Optional correction: what the formula actually assumes

Define pass as the positive label. Sensitivity Se is the fraction of human-passed answers that the judge passes; specificity Sp is the fraction of human-failed answers that the judge fails. If these rates transfer to the assessed population, the observed judge-pass fraction q relates to the true pass fraction p:

q = Se × p + (1 − Sp) × (1 − p)
p = (q + Sp − 1) / (Se + Sp − 1)

Se = 0.90, Sp = 0.80, q = 0.76
p = (0.76 + 0.80 − 1) / 0.70 = 0.80

This is the Rogan–Gladen correction, used by judgy. Precision and recall alone are not interchangeable with sensitivity and specificity. A denominator near zero is unstable; an estimate outside 0–1 signals a problem to investigate. Propagate both calibration and evaluation uncertainty, including shared baseline/candidate calibration. Changing answer style can change judge errors, breaking transfer. The point estimate above supplies no confidence interval by itself.

Prediction-powered inference is a different framework that combines labeled evidence and model predictions for statistical inference. It is not another name for this binary correction. Neither method guarantees that every good release passes or every bad release is blocked.

7. Make the statistical decision explicit

Define quality delta as candidate pass rate − baseline pass rate on the same eligible cases. A non-inferiority margin is the largest tolerated decline for a particular metric. Here use two percentage points (0.02) as a product decision, separate from hard requirements.

Evidence for delta Decision under this illustrative rule Meaning
95% interval −1.2 to +0.6 points Pass this quality criterion Lower bound remains above −2 points
95% interval −2.8 to +0.3 points Hold Data permit both acceptable and unacceptable change
95% interval −4.0 to −2.5 points Block Entire interval violates the margin
Any confirmed critical disclosure Block Hard requirement takes precedence
Missing cases, expired calibration or stale build Hold Required evidence is absent or not applicable

A confidence interval describes the long-run coverage of a statistical procedure under its assumptions; it is not a guarantee about a release. A paired bootstrap resamples old/new results together because both answers came from the same case. If cases share contracts or conversations, resample those groups together and preserve stratification/weights. Repeated generations are nested within cases, not additional independent users. See confidence interval principles.

Why there is no universal 1,200-case minimum

Suppose 1,200 independent paired cases produce 1,080 passes in both builds, 30 baseline-only passes, 18 candidate-only passes and 72 failures in both. The delta is (18 − 30)/1,200 = −1%. For a rough normal approximation to the mean paired difference:

Discordant fraction = (30 + 18) / 1,200 = 0.04
Standard error ≈ sqrt((0.04 − 0.01²) / 1,200) = 0.00577
Approximate 95% interval ≈ −1% ± 1.13 percentage points

The lower bound is about −2.13%, so this illustration is inconclusive at a two-point margin. At 4,000 independent cases with the same observed proportions, the approximate interval narrows to −1% ± 0.62 points. This illustrates dependence on variance and sample size; it is not a power calculation, a prescription for small/rare-event samples, or a correction for a biased judge. Plan power and resampling around the actual metric, dependence and tolerated risk.

Predeclare required slices, decision rules and when additional evidence may be collected. Repeatedly rerunning until green creates optional-stopping bias. Many separate tests can also generate false alarms or selected apparent gains. Specify the family of claims and use a suitable simultaneous or sequential procedure where required. Requiring all predefined non-inferiority criteria to pass has different statistical logic from selecting whichever metric improves; do not apply one generic multiple-testing shortcut to both.

Executable decision policy

This function consumes validated, trusted evaluator results. It classifies evidence; it does not compute intervals, authenticate workers or replace transactional publication of a required check. Bounds and margins use fractions, not percentage-point numbers.

import math

def eval_gate(*, expected_manifest, observed_manifest, complete,
              hard_failure, calibrated, required_margins, intervals):
    if not expected_manifest or observed_manifest != expected_manifest:
        return "hold_stale_evidence"
    if hard_failure is True:
        return "block_critical_failure"
    if hard_failure is not False or complete is not True or calibrated is not True:
        return "hold_incomplete_evidence"
    if not required_margins or set(intervals) != set(required_margins):
        return "hold_missing_criteria"
    statuses = []
    for criterion, margin in required_margins.items():
        bounds = intervals[criterion]
        if not isinstance(bounds, (list, tuple)) or len(bounds) != 2:
            return "hold_invalid_interval"
        low, high = bounds
        values = [margin, low, high]
        if any(type(v) not in (int, float) or not -1 <= v <= 1
               or not math.isfinite(v) for v in values):
            return "hold_invalid_interval"
        if margin < 0 or low > high:
            return "hold_invalid_interval"
        statuses.append("pass" if low >= -margin else
                        "block" if high < -margin else "hold")
    if "block" in statuses:
        return "block_quality_regression"
    if "hold" in statuses:
        return "hold_more_evidence"
    return "pass"

The full manifest includes judge, datasets and policies as well as build identity. A trusted publisher checks the decision's signature, expiry and current revision before setting CI status. The exact margin and boundary convention belong in the versioned policy.

8. Cache reusable work and promote the tested artifact

Cache generation by all behavior-affecting dependencies: code artifact, input, prompts, model/configuration, retrieval snapshot, tool fixtures, policy, tenant scope and trial ID. Cache grading separately by exact output, expected properties, rubric and judge configuration. A prompt-only key misses orchestration changes. Reusing one cached answer for every trial falsely suggests deterministic behavior.

Replaying fixtures tests controlled behavior; live integration tests verify that the real connector still works. When intentionally comparing retrieval versions, keep each build's retrieval behavior but use compatible input corpora and rights snapshots. Do not accidentally erase the component difference being evaluated.

Architecture / visual model
sequenceDiagram participant CI as Trusted CI controller participant W as Isolated workers participant G as Gate and evidence store participant Q as Merge queue participant D as Deployment controller CI->>W: Exact candidate and baseline manifests W->>G: Completed cases, outputs and grader records G-->>CI: Pass, block or hold with evidence digest alt Required evidence passes CI->>Q: Successful check on merge-group revision Q-->>D: Approved immutable artifact manifest D->>D: Verify evidence identity and compatibility D->>D: Isolated shadow then bounded canary alt Live release criteria pass D->>D: Increase exposure else Critical or material regression D->>D: Stop exposure and restore compatible release end else Failure or uncertainty CI->>Q: Required check remains unsatisfied end
Read diagram source
sequenceDiagram
    participant CI as Trusted CI controller
    participant W as Isolated workers
    participant G as Gate and evidence store
    participant Q as Merge queue
    participant D as Deployment controller
    CI->>W: Exact candidate and baseline manifests
    W->>G: Completed cases, outputs and grader records
    G-->>CI: Pass, block or hold with evidence digest
    alt Required evidence passes
        CI->>Q: Successful check on merge-group revision
        Q-->>D: Approved immutable artifact manifest
        D->>D: Verify evidence identity and compatibility
        D->>D: Isolated shadow then bounded canary
        alt Live release criteria pass
            D->>D: Increase exposure
        else Critical or material regression
            D->>D: Stop exposure and restore compatible release
        end
    else Failure or uncertainty
        CI->>Q: Required check remains unsatisfied
    end

Shadow execution must suppress real writes, duplicate emails and external side effects. A canary exposes a bounded population and must last long enough for relevant outcomes, including recontact or delayed tool failures. Do not require one arbitrary 30-minute duration for every product.

Rollback restores a compatible code, prompt, model, tools, policy and index combination. Keep old schemas readable or provide a tested migration strategy. A rollback cannot undo a payment already made, delete an email already delivered or restore a vendor model that has been retired. Include compensating actions and a prevalidated degraded mode.

9. Failure modes and repairs

Failure Diagnosis and repair Cost or remaining limitation
F1: Judge drift Compare with fresh expert labels, freeze affected promotions and regrade both builds Calibration work; pinned models can still be retired
F2: Eval overfitting Maintain fresh restricted assessment and detect training/prompt overlap New labels; hidden outcomes still leak through repeated tuning
F3: Sample misses a domain regression Require relevant slices and adequate evidence before merge Larger suite; targeted cases do not estimate prevalence
F4: Full-run spend exceeds allowance Reserve estimated cost before starting, bound retries and require extra budget approval Delayed run; budget exhaustion is not a pass
F5: High questionable-block rate Inspect flaky tests, inconsistent labels, grader changes and actual regressions Reviewer work; do not tune thresholds to a desired block percentage
F6: Holdout exposure Quarantine exposed cases, invalidate contaminated claims and create fresh independent cases Reassessment; hashing reports is not complete secrecy
F7: Judge deprecation Calibrate a replacement and rerun comparable baselines before retirement Temporary double evaluation; no universal notice period
F8: Worker or quota saturation Cancel obsolete jobs, enforce priority/fairness and add approved capacity Required checks remain pending; no silent sample reduction

Also test a spoofed result, incomplete manifest, stale merge revision, unauthorized trace read and output that tries to instruct the judge to award a pass. Judge prompts delimit evidence as data; execution isolation and trusted reporting enforce the security boundary.

10. Operational Considerations and economics

Measurements and runbooks

Signal Denominator or meaning Action
Evaluation completion p95 Admission to complete required evidence Split queue, provider and scorer bottlenecks
Block rate Blocked runs / completed assessed runs Diagnostic only; includes true and false blocks
Missing-case rate Missing required observations / scheduled observations Hold affected decision and repair run
False-pass/false-block estimates Expert-audited decisions, with sampling design Revisit grader and policy evidence
Escape rate Material production regressions under a defined exposure denominator Investigate each critical incident
Cost Actual billed tokens/tools plus runner and reviewer costs Compare with reservation and cap retries

Every report leads with the decision, affected revision and reason. Show newly failing/passing examples, severity, sample/cluster counts, intervals, coverage gaps, baseline/grade versions and cost. Reveal development examples only to authorized engineers. Restricted assessments need an independent process to adjudicate failures without gradually publishing the holdout.

Worked cost model

Use Sonnet 5 standard text prices of $2/M input and $10/M output. Assume both a generation call and a grading call contain 4,000 input and 600 output tokens: each costs 0.008 + 0.006 = $0.014. This is a billing assumption, not a required prompt length. Claude pricing.

Model work Calculation Cost
Default PR, one criterion 500 cases × 2 builds × 2 calls × $0.014 $28
Broader comparison 4,000 × 2 × 2 × $0.014 $224
Single-build nightly 4,000 × 1 × 2 × $0.014 $112
Weekly planned model spend 50 × $28 + 8 additional broader runs × $224 + 7 nightly runs × $112 $3,976
Monthly model allowance Weekly × 52 / 12, then 10% retry/repetition allowance $18,952.27
Runner, storage and telemetry Explicit planning allowance $1,000/month
Expert calibration and failure triage 20 hours/week × $120/hour × 52 / 12 $10,400/month
Included operating total Models + platform + expert time $30,352.27/month

This includes baseline generation; compatible cached outputs can reduce cost, while extra criteria, longer answers and protected-branch runs increase it. One more separately called criterion adds 500 × 2 × $0.014 = $14 to a default PR, already exceeding the $40 target when added to $28. Reusing one output avoids regenerating it, but grading still costs money. Combine compatible criteria in one validated rubric or request budget; never silently drop a critical gate.

A small lost-renewal probability can justify significant evaluation spend, but avoid claiming a particular customer loss was prevented without evidence. At this scenario's $364,227 yearly included cost, a hypothetical $4M avoided loss needs about a 9.1-percentage-point reduction in its annual probability to break even if that is the only benefit. Engineering implementation, ordinary code CI and other business costs remain outside this estimate.

Quarterly, review taxonomy, permissions/retention, escaped failures, calibration, budget and exceptions. The evidence pack records methodology, suite versions, sampled decisions, independent labels, unresolved limitations and signoff. It supports a review; it does not itself establish regulatory compliance.

Interview follow-ups

Q1. Why compare on the same cases instead of comparing two pass rates from different runs?

The same cases control for task difficulty and permit analysis of which outcomes changed. Both builds also need compatible fixtures and the same grader. Pairing does not remove model randomness, judge bias or correlated contract examples; the trial and resampling design must address those separately.

Q2. The quality interval crosses the tolerated decline. Is that a failure?

It is insufficient evidence for non-inferiority under this rule. Hold the change, diagnose the affected cases or collect the prespecified additional evidence. Distinguish that uncertainty from a clearly unacceptable regression. Repeatedly rerunning until the lower bound happens to pass changes the statistical procedure.

Q3. Can 99% overall accuracy justify one tenant-data leak?

No. Tenant isolation is a hard invariant in this design. Block, repair the access boundary and expand the relevant security tests. Averaging that incident into a quality score would erase the actual release policy. A sampled test suite still cannot prove zero future leaks.

Q4. Why not solve judge bias with a correction library?

A correction estimates a defined quantity under calibration assumptions. It cannot fix a poor rubric, contaminated labels, changed error rates or an unauthorized tool action. I would validate the judge and uncertainty method, inspect important slices and keep independently executable requirements separate.

Q5. The queue is full and the release is urgent. What changes?

Cancel superseded work, prioritize the release and use available approved capacity. Preserve required evidence. An authorized exception must name the omitted evidence, exposure bound, mitigations and expiry; it must not masquerade as a green evaluation. Some security requirements remain non-waivable under the product's policy.

Q6. Two PRs each passed but fail when merged together. How do you prevent this?

Evaluate the prospective merged revision and bind results to its full behavior manifest. Re-run relevant checks when the merge group changes. Promote that tested artifact; rebuilding with unpinned dependencies can create another untested version even from the same source revision.

Q7. What makes this a strong manager-level design?

Give each decision an accountable owner: domain experts define correct outcomes, engineering owns reproducibility, security owns access/abuse requirements with the team, and the release owner controls exposure. Track user outcomes and triage time, not just green percentages. Budget expert calibration as part of the system rather than assuming automated grading eliminates it.

60-second interview answer

I would build a trusted evaluation pipeline around the exact candidate and baseline manifests. Fast contract tests run first, followed by relevant quality and severe-risk suites. Semantic judges need expert calibration, while paired comparisons need a declared uncertainty and decision policy. The gate passes, blocks or holds for more evidence; missing checks never become a pass. I would validate the prospective merged revision, promote the tested artifact through bounded exposure and retain a compatible rollback. Reports, budgets and named owners make the process usable under delivery pressure.

Remember: Define risks → Version evidence → Compare paired cases → Decide explicitly → Release the tested artifact.

Final notes: Keep hard requirements separate from average quality. Distinguish representative estimates from targeted tests. Price both builds and human review. A green check is only useful when it applies to the release actually shipped.

Related concepts: CI/CD for AI, LLM evaluation, guardrails, multi-tenant training.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Case Study: Customer-Specific Fine-Tuning Platform
NEXT LESSONDesign a Customer-Specific Distillation Pipeline →

Explore the diagram