Hypothetical interview scenario. Workloads, tolerances, costs and service targets are assumptions for this design. They are not Learnastra operating results or universal standards.
Interview focus: Decide whether a specific AI change is ready to merge and release, using reproducible evidence and an explicit treatment of uncertainty.
1. Define the problem and scope
Continuous integration (CI) automatically builds and tests proposed software changes. Continuous delivery keeps validated changes ready for release; continuous deployment also releases them automatically after the required checks. An evaluation gate is a required decision based on tests of the AI system's behavior. It complements ordinary software tests.
Design the gate for a 28-engineer team with four product teams and 50 AI-related pull requests (PRs) a week. The product answers questions about customer contracts using retrieval and tools. A shorter-answer prompt can omit a renewal exception while making the average style score improve. The gate must detect that consequential omission, explain the evidence and identify the exact build it assessed.
Clarify with the interviewer: Which failures must prevent release? Who labels correct contract interpretations? Does this gate allow merge, production exposure, or both? Can evaluation access customer data? What is the acceptable delay when evidence is inconclusive?
Functional requirements
- Test candidate and baseline behavior on versioned inputs, source documents, tool fixtures and scoring criteria.
- Run deterministic contract checks, representative quality comparisons and targeted severe-failure tests as separate suites.
- Select additional coverage from the changed components and affected customer domains.
- Validate semantic judges against expert labels and retain their calibration versions.
- Report pass, block or hold for more evidence, with changed examples and per-criterion results.
- Bind the decision to a build and evidence manifest; invalidate it when a relevant dependency changes.
- Promote an approved artifact through controlled deployment and retain a compatible rollback release.
- Maintain private traces, reviewer decisions, budgets and an auditable exception process.
Non-functional requirements
- Integrity: no merge based on a stale, incomplete, cancelled or untrusted check. A known critical tenant-data disclosure blocks regardless of aggregate quality.
- Latency: aim for p95 evaluation completion within 60 minutes of admission, leaving review time within a 90-minute PR-to-merge target. Report queue and execution time separately.
- Cost: default PR model spend below $40; explicit approval above that. A $1,200 full-run ceiling is an admission limit, not the predicted price.
- Privacy: tenant-approved data, minimum runner privileges and restricted access to customer examples and independent holdouts.
- Reproducibility: preserve the tested configuration and evidence. Model nondeterminism still requires measured repeated trials.
- Operations: failures remain visible and pending until resolved; quarterly methodology review is this scenario's governance requirement, not a universal legal mandate.
Out of scope: proving that no future failure can occur, letting a model approve its own security exceptions, and executing irreversible customer actions in test environments.
2. Establish a useful baseline
Begin with code tests plus a small, reviewed regression suite for known problems: correct currency, permitted tenant, valid tool arguments, real citation IDs and the renewal exception. Store expected properties rather than an exact prose answer. Run the candidate in an isolated environment, compare against the current production release and require a reviewer to inspect failures.
| Baseline weakness | Why it matters | Next improvement and cost |
|---|---|---|
| Only known examples | Misses new tasks and languages | Add a representative sample and fresh error analysis; ongoing labeling work |
| One overall judge score | Style can mask omitted obligations | Separate critical requirements and quality dimensions; more reporting |
| Saved old score versus newly graded candidate | A judge change looks like a product change | Grade comparable outputs with the same rubric/version; extra calls |
| PR passes before another PR merges | Their combined behavior is untested | Evaluate the prospective merge revision; extra queue work |
| Green offline report means immediate broad release | Fixtures miss external-service and live-traffic behavior | Integration tests and canary exposure; slower rollout |
A small suite is useful development feedback. Calling it evidence of population-wide quality requires a justified sampling plan.
3. Size the evaluation workload
Use a 4,000-case broader quality suite and a 500-case default PR suite as initial planning sizes, not statistical guarantees. Mandatory severe-risk and previously broken behaviors are separately identified; repeated or deliberately difficult examples must not be counted as independent representative observations.
| Quantity | Calculation | Implication |
|---|---|---|
| Default work | 500 cases × 2 builds × generation and grading | 2,000 model calls per PR for one criterion |
| Default week | 50 × 2,000 | 100,000 calls before full runs/retries |
| Sequential model time | 2,000 × assumed 3 seconds | 6,000 seconds of call time |
| Ideal 20-call concurrency | 6,000 / 20 | 5 minutes lower-bound service time; dependencies, rate limits and long tails add time |
| Weekly execution demand | 50 × assumed 12 runner-minutes | 10 runner-hours; averages do not size a release-day burst |
| Ten simultaneous PRs | 10 × 20 active calls | Up to 200 calls competing for provider limits |
Admit jobs against both request and token quotas. At 4,000 input and 600 output tokens, 20 calls each taking three seconds would demand roughly 1.84M combined tokens/minute. Providers may meter input and output separately. More workers cannot bypass those limits. Reserve capacity for the required merge checks, cancel superseded runs and cap retries.
4. Detailed architecture and evidence contracts
Read diagram source
flowchart TB
PR[PR and prospective merge revision] --> PLAN[Trusted planner<br/>change scope and required suites]
PLAN --> ADMIT[Budget and rate-limit admission]
DATA[Versioned cases<br/>source snapshots and tool fixtures] --> ADMIT
ADMIT --> RUN[Isolated baseline and candidate workers]
RUN --> EXACT[Exact contract and security assertions]
RUN --> SEM[Semantic graders<br/>pinned rubric and model]
HUMAN[Independent expert labels<br/>calibration and disagreements] --> SEM
EXACT --> EVIDENCE[Append-only evidence manifests]
SEM --> EVIDENCE
EVIDENCE --> GATE[Trusted gate<br/>identity, coverage and uncertainty]
GATE -->|Block or hold| REPORT[Scoped report to developer]
GATE -->|Pass| CHECK[Required check on exact revision]
CHECK --> RELEASE[Artifact promotion<br/>shadow and canary stages]
RELEASE --> MON[Live outcomes and rollback controller]
MON --> DATA
The planner, grader and check reporter run from a trusted version. A PR that can replace its own gate with return pass defeats the design. Separate untrusted candidate execution from the authority to read restricted holdouts or write a successful check. Use short-lived, scoped credentials, isolated runners and redacted artifacts. Do not expose deployment secrets to untrusted PR code or treat log masking as complete data protection. GitHub Actions secure use.
Practical implementation choices
| Component | Starting choice | Why and when to change |
|---|---|---|
| CI control | GitHub Actions plus protected required checks | Fits the assumed PR workflow; keep the gate portable to other CI systems |
| Exact assertions | Pytest and domain-specific validators | Transparent expected outcomes; do not encode semantic truth as fragile string equality |
| Case and result storage | Versioned JSONL in approved object storage plus a relational run catalog | Cheap immutable evidence and queryable status; catalog and blobs need consistent publication |
| Traces and comparisons | Langfuse with access controls and retention | Supports inspecting evaluation experiments; storing a score does not enforce the release policy |
| Human annotation | Restricted review UI or Argilla | Independent labeling and disagreement resolution; budget domain-expert time |
Langfuse evaluation documentation distinguishes offline experiments from online evaluation. Implement the required-check policy explicitly instead of assuming an observability dashboard blocks deployment.
| Record | Required content | Contract |
|---|---|---|
| Case | ID/version, input, expected properties, source permissions, slice, sampling weight and cluster ID | Source and label versions are immutable for a run |
| Build manifest | Code/artifact digest, prompts, model/config, index snapshot, tools and policies | Describes the executed behavior, not only a Git SHA |
| Evaluation run | Run ID, exact baseline/candidate manifests, suite, judge, repetitions, budgets and status | Deduplicate triggers by the full run specification |
| Observation | Case/build/trial IDs, output, assertion results, judge details, duration and billable usage | Missing output is recorded, not silently excluded |
| Gate decision | Required checks, coverage, per-slice deltas/intervals, policy version and decision | Only trusted evaluator evidence can produce a pass |
| Exception | Owner, reason, impacted criteria, exposure limit, compensating controls and expiry | Explicitly distinguished from an ordinary pass |
| Release | Artifact manifest, approved evidence digest, rollout state and prior compatible release | Promotion checks identity again |
A possible internal API is POST /eval-runs with immutable manifest IDs and an idempotency key. GET /eval-runs/{id} exposes queued, running, completed, failed or cancelled; completed evaluation is not synonymous with a passing gate. POST /release-decisions accepts trusted evidence references, not a candidate-supplied score.
When using GitHub's merge queue, run required checks on the merge group revision through its merge_group event. It includes the current base and preceding queued changes, so it differs from the original PR revision. A green check from the old revision is insufficient. Apply the equivalent contract in another CI provider. GitHub merge queues.
5. Build the right test sets and rubric
Separate three kinds of evidence
| Suite | How cases are selected | What its score can tell you |
|---|---|---|
| Representative quality | Permitted traffic sampled by a documented design | Estimated performance for that population, using correct weights |
| Targeted regressions and attacks | Known failures, synthetic attacks and important edge cases | Whether these specified risks still fail; not their population frequency |
| Restricted assessment | Fresh independent cases, access controlled and not repeatedly tuned against | A less contaminated assessment of the selected candidate |
Stratify by contract type, language and workflow where needed. If a rare category is oversampled, weight it back for the overall population estimate and still show its separate result. One example per category gives coverage of names, not reliable measurement. With a true independent 1% failure rate, ten random cases miss every failure with probability 0.99^10 ≈ 90.4%.
Keep serious historical regressions required while they remain applicable. Archive obsolete behaviors with a reason rather than deleting inconvenient failures. Rotate fresh assessment cases based on drift and exposure; “replace 10% quarterly” does not itself prevent overfitting. Repeated pass/fail feedback also leaks information about a holdout, even if raw examples are hidden.
Turn observed errors into assertions
Open coding means labeling observed failures with specific descriptions. Axial coding groups and relates those labels into useful families. Two reviewers apply the definitions to fresh examples and resolve ambiguity before scaling annotation.
| Observed failure | Failure family | Test and owner |
|---|---|---|
| Current contract retrieved, old notice period used | Wrong source applicability | Check date and contract edition; retrieval/domain team |
| Renewal exception omitted | Incomplete answer | Rubric requires exception when condition applies; domain reviewers |
| Citation ID does not exist | Invented evidence | Exact ID membership plus semantic support check; application team |
| Refund currency differs from order | Invalid action | Exact pre-execution currency assertion; payments integration team |
| Correct answer contains another tenant's clause | Unauthorized disclosure | Isolated tenant fixtures and access tests; security and application teams |
A citation ID resolving is necessary but does not establish that the cited text supports the claim. Likewise, valid JSON can still contain a wrong amount. Define refusal errors in both directions: failing to refuse a prohibited request and refusing a permitted one.
6. Check the judge before trusting its scores
An LLM judge grades an output against a rubric. It is a noisy measurement component, not ground truth. Use code where the result is objectively executable; use domain experts where correctness requires interpretation, and test whether the judge agrees for each relevant criterion.
- Create written labels with examples and counterexamples. Resolve human disagreement and retain uncertainty.
- Keep judge-development, selection and final assessment data distinct. A 60/20/20 split is an option, not a requirement; all splits need enough relevant examples.
- Pin grader prompt, model, rubric and generation settings. Randomize answer order in pairwise evaluation and test for position, verbosity and self-preference bias.
- Track false passes and false failures by domain. Overall accuracy and Cohen's kappa do not establish adequate recall for rare harmful errors.
- Recalibrate when the application, judge or population changes. Monthly fresh labeling is a cadence hypothesis; a material change requires immediate review.
Position and verbosity biases are documented in the original LLM-as-a-judge study. Two judges agreeing is useful diagnostic evidence, but they may share the same mistake. A larger model can help adjudicate cases without replacing expert labels.
Optional correction: what the formula actually assumes
Define pass as the positive label. Sensitivity Se is the fraction of human-passed answers that the judge passes; specificity Sp is the fraction of human-failed answers that the judge fails. If these rates transfer to the assessed population, the observed judge-pass fraction q relates to the true pass fraction p:
q = Se × p + (1 − Sp) × (1 − p)
p = (q + Sp − 1) / (Se + Sp − 1)
Se = 0.90, Sp = 0.80, q = 0.76
p = (0.76 + 0.80 − 1) / 0.70 = 0.80
This is the Rogan–Gladen correction, used by judgy. Precision and recall alone are not interchangeable with sensitivity and specificity. A denominator near zero is unstable; an estimate outside 0–1 signals a problem to investigate. Propagate both calibration and evaluation uncertainty, including shared baseline/candidate calibration. Changing answer style can change judge errors, breaking transfer. The point estimate above supplies no confidence interval by itself.
Prediction-powered inference is a different framework that combines labeled evidence and model predictions for statistical inference. It is not another name for this binary correction. Neither method guarantees that every good release passes or every bad release is blocked.
7. Make the statistical decision explicit
Define quality delta as candidate pass rate − baseline pass rate on the same eligible cases. A non-inferiority margin is the largest tolerated decline for a particular metric. Here use two percentage points (0.02) as a product decision, separate from hard requirements.
| Evidence for delta | Decision under this illustrative rule | Meaning |
|---|---|---|
| 95% interval −1.2 to +0.6 points | Pass this quality criterion | Lower bound remains above −2 points |
| 95% interval −2.8 to +0.3 points | Hold | Data permit both acceptable and unacceptable change |
| 95% interval −4.0 to −2.5 points | Block | Entire interval violates the margin |
| Any confirmed critical disclosure | Block | Hard requirement takes precedence |
| Missing cases, expired calibration or stale build | Hold | Required evidence is absent or not applicable |
A confidence interval describes the long-run coverage of a statistical procedure under its assumptions; it is not a guarantee about a release. A paired bootstrap resamples old/new results together because both answers came from the same case. If cases share contracts or conversations, resample those groups together and preserve stratification/weights. Repeated generations are nested within cases, not additional independent users. See confidence interval principles.
Why there is no universal 1,200-case minimum
Suppose 1,200 independent paired cases produce 1,080 passes in both builds, 30 baseline-only passes, 18 candidate-only passes and 72 failures in both. The delta is (18 − 30)/1,200 = −1%. For a rough normal approximation to the mean paired difference:
Discordant fraction = (30 + 18) / 1,200 = 0.04
Standard error ≈ sqrt((0.04 − 0.01²) / 1,200) = 0.00577
Approximate 95% interval ≈ −1% ± 1.13 percentage points
The lower bound is about −2.13%, so this illustration is inconclusive at a two-point margin. At 4,000 independent cases with the same observed proportions, the approximate interval narrows to −1% ± 0.62 points. This illustrates dependence on variance and sample size; it is not a power calculation, a prescription for small/rare-event samples, or a correction for a biased judge. Plan power and resampling around the actual metric, dependence and tolerated risk.
Predeclare required slices, decision rules and when additional evidence may be collected. Repeatedly rerunning until green creates optional-stopping bias. Many separate tests can also generate false alarms or selected apparent gains. Specify the family of claims and use a suitable simultaneous or sequential procedure where required. Requiring all predefined non-inferiority criteria to pass has different statistical logic from selecting whichever metric improves; do not apply one generic multiple-testing shortcut to both.
Executable decision policy
This function consumes validated, trusted evaluator results. It classifies evidence; it does not compute intervals, authenticate workers or replace transactional publication of a required check. Bounds and margins use fractions, not percentage-point numbers.
import math
def eval_gate(*, expected_manifest, observed_manifest, complete,
hard_failure, calibrated, required_margins, intervals):
if not expected_manifest or observed_manifest != expected_manifest:
return "hold_stale_evidence"
if hard_failure is True:
return "block_critical_failure"
if hard_failure is not False or complete is not True or calibrated is not True:
return "hold_incomplete_evidence"
if not required_margins or set(intervals) != set(required_margins):
return "hold_missing_criteria"
statuses = []
for criterion, margin in required_margins.items():
bounds = intervals[criterion]
if not isinstance(bounds, (list, tuple)) or len(bounds) != 2:
return "hold_invalid_interval"
low, high = bounds
values = [margin, low, high]
if any(type(v) not in (int, float) or not -1 <= v <= 1
or not math.isfinite(v) for v in values):
return "hold_invalid_interval"
if margin < 0 or low > high:
return "hold_invalid_interval"
statuses.append("pass" if low >= -margin else
"block" if high < -margin else "hold")
if "block" in statuses:
return "block_quality_regression"
if "hold" in statuses:
return "hold_more_evidence"
return "pass"
The full manifest includes judge, datasets and policies as well as build identity. A trusted publisher checks the decision's signature, expiry and current revision before setting CI status. The exact margin and boundary convention belong in the versioned policy.
8. Cache reusable work and promote the tested artifact
Cache generation by all behavior-affecting dependencies: code artifact, input, prompts, model/configuration, retrieval snapshot, tool fixtures, policy, tenant scope and trial ID. Cache grading separately by exact output, expected properties, rubric and judge configuration. A prompt-only key misses orchestration changes. Reusing one cached answer for every trial falsely suggests deterministic behavior.
Replaying fixtures tests controlled behavior; live integration tests verify that the real connector still works. When intentionally comparing retrieval versions, keep each build's retrieval behavior but use compatible input corpora and rights snapshots. Do not accidentally erase the component difference being evaluated.
Read diagram source
sequenceDiagram
participant CI as Trusted CI controller
participant W as Isolated workers
participant G as Gate and evidence store
participant Q as Merge queue
participant D as Deployment controller
CI->>W: Exact candidate and baseline manifests
W->>G: Completed cases, outputs and grader records
G-->>CI: Pass, block or hold with evidence digest
alt Required evidence passes
CI->>Q: Successful check on merge-group revision
Q-->>D: Approved immutable artifact manifest
D->>D: Verify evidence identity and compatibility
D->>D: Isolated shadow then bounded canary
alt Live release criteria pass
D->>D: Increase exposure
else Critical or material regression
D->>D: Stop exposure and restore compatible release
end
else Failure or uncertainty
CI->>Q: Required check remains unsatisfied
end
Shadow execution must suppress real writes, duplicate emails and external side effects. A canary exposes a bounded population and must last long enough for relevant outcomes, including recontact or delayed tool failures. Do not require one arbitrary 30-minute duration for every product.
Rollback restores a compatible code, prompt, model, tools, policy and index combination. Keep old schemas readable or provide a tested migration strategy. A rollback cannot undo a payment already made, delete an email already delivered or restore a vendor model that has been retired. Include compensating actions and a prevalidated degraded mode.
9. Failure modes and repairs
| Failure | Diagnosis and repair | Cost or remaining limitation |
|---|---|---|
| F1: Judge drift | Compare with fresh expert labels, freeze affected promotions and regrade both builds | Calibration work; pinned models can still be retired |
| F2: Eval overfitting | Maintain fresh restricted assessment and detect training/prompt overlap | New labels; hidden outcomes still leak through repeated tuning |
| F3: Sample misses a domain regression | Require relevant slices and adequate evidence before merge | Larger suite; targeted cases do not estimate prevalence |
| F4: Full-run spend exceeds allowance | Reserve estimated cost before starting, bound retries and require extra budget approval | Delayed run; budget exhaustion is not a pass |
| F5: High questionable-block rate | Inspect flaky tests, inconsistent labels, grader changes and actual regressions | Reviewer work; do not tune thresholds to a desired block percentage |
| F6: Holdout exposure | Quarantine exposed cases, invalidate contaminated claims and create fresh independent cases | Reassessment; hashing reports is not complete secrecy |
| F7: Judge deprecation | Calibrate a replacement and rerun comparable baselines before retirement | Temporary double evaluation; no universal notice period |
| F8: Worker or quota saturation | Cancel obsolete jobs, enforce priority/fairness and add approved capacity | Required checks remain pending; no silent sample reduction |
Also test a spoofed result, incomplete manifest, stale merge revision, unauthorized trace read and output that tries to instruct the judge to award a pass. Judge prompts delimit evidence as data; execution isolation and trusted reporting enforce the security boundary.
10. Operational Considerations and economics
Measurements and runbooks
| Signal | Denominator or meaning | Action |
|---|---|---|
| Evaluation completion p95 | Admission to complete required evidence | Split queue, provider and scorer bottlenecks |
| Block rate | Blocked runs / completed assessed runs | Diagnostic only; includes true and false blocks |
| Missing-case rate | Missing required observations / scheduled observations | Hold affected decision and repair run |
| False-pass/false-block estimates | Expert-audited decisions, with sampling design | Revisit grader and policy evidence |
| Escape rate | Material production regressions under a defined exposure denominator | Investigate each critical incident |
| Cost | Actual billed tokens/tools plus runner and reviewer costs | Compare with reservation and cap retries |
Every report leads with the decision, affected revision and reason. Show newly failing/passing examples, severity, sample/cluster counts, intervals, coverage gaps, baseline/grade versions and cost. Reveal development examples only to authorized engineers. Restricted assessments need an independent process to adjudicate failures without gradually publishing the holdout.
Worked cost model
Use Sonnet 5 standard text prices of $2/M input and $10/M output. Assume both a generation call and a grading call contain 4,000 input and 600 output tokens: each costs 0.008 + 0.006 = $0.014. This is a billing assumption, not a required prompt length. Claude pricing.
| Model work | Calculation | Cost |
|---|---|---|
| Default PR, one criterion | 500 cases × 2 builds × 2 calls × $0.014 | $28 |
| Broader comparison | 4,000 × 2 × 2 × $0.014 | $224 |
| Single-build nightly | 4,000 × 1 × 2 × $0.014 | $112 |
| Weekly planned model spend | 50 × $28 + 8 additional broader runs × $224 + 7 nightly runs × $112 | $3,976 |
| Monthly model allowance | Weekly × 52 / 12, then 10% retry/repetition allowance | $18,952.27 |
| Runner, storage and telemetry | Explicit planning allowance | $1,000/month |
| Expert calibration and failure triage | 20 hours/week × $120/hour × 52 / 12 | $10,400/month |
| Included operating total | Models + platform + expert time | $30,352.27/month |
This includes baseline generation; compatible cached outputs can reduce cost, while extra criteria, longer answers and protected-branch runs increase it. One more separately called criterion adds 500 × 2 × $0.014 = $14 to a default PR, already exceeding the $40 target when added to $28. Reusing one output avoids regenerating it, but grading still costs money. Combine compatible criteria in one validated rubric or request budget; never silently drop a critical gate.
A small lost-renewal probability can justify significant evaluation spend, but avoid claiming a particular customer loss was prevented without evidence. At this scenario's $364,227 yearly included cost, a hypothetical $4M avoided loss needs about a 9.1-percentage-point reduction in its annual probability to break even if that is the only benefit. Engineering implementation, ordinary code CI and other business costs remain outside this estimate.
Quarterly, review taxonomy, permissions/retention, escaped failures, calibration, budget and exceptions. The evidence pack records methodology, suite versions, sampled decisions, independent labels, unresolved limitations and signoff. It supports a review; it does not itself establish regulatory compliance.
Interview follow-ups
Q1. Why compare on the same cases instead of comparing two pass rates from different runs?
The same cases control for task difficulty and permit analysis of which outcomes changed. Both builds also need compatible fixtures and the same grader. Pairing does not remove model randomness, judge bias or correlated contract examples; the trial and resampling design must address those separately.
Q2. The quality interval crosses the tolerated decline. Is that a failure?
It is insufficient evidence for non-inferiority under this rule. Hold the change, diagnose the affected cases or collect the prespecified additional evidence. Distinguish that uncertainty from a clearly unacceptable regression. Repeatedly rerunning until the lower bound happens to pass changes the statistical procedure.
Q3. Can 99% overall accuracy justify one tenant-data leak?
No. Tenant isolation is a hard invariant in this design. Block, repair the access boundary and expand the relevant security tests. Averaging that incident into a quality score would erase the actual release policy. A sampled test suite still cannot prove zero future leaks.
Q4. Why not solve judge bias with a correction library?
A correction estimates a defined quantity under calibration assumptions. It cannot fix a poor rubric, contaminated labels, changed error rates or an unauthorized tool action. I would validate the judge and uncertainty method, inspect important slices and keep independently executable requirements separate.
Q5. The queue is full and the release is urgent. What changes?
Cancel superseded work, prioritize the release and use available approved capacity. Preserve required evidence. An authorized exception must name the omitted evidence, exposure bound, mitigations and expiry; it must not masquerade as a green evaluation. Some security requirements remain non-waivable under the product's policy.
Q6. Two PRs each passed but fail when merged together. How do you prevent this?
Evaluate the prospective merged revision and bind results to its full behavior manifest. Re-run relevant checks when the merge group changes. Promote that tested artifact; rebuilding with unpinned dependencies can create another untested version even from the same source revision.
Q7. What makes this a strong manager-level design?
Give each decision an accountable owner: domain experts define correct outcomes, engineering owns reproducibility, security owns access/abuse requirements with the team, and the release owner controls exposure. Track user outcomes and triage time, not just green percentages. Budget expert calibration as part of the system rather than assuming automated grading eliminates it.
60-second interview answer
I would build a trusted evaluation pipeline around the exact candidate and baseline manifests. Fast contract tests run first, followed by relevant quality and severe-risk suites. Semantic judges need expert calibration, while paired comparisons need a declared uncertainty and decision policy. The gate passes, blocks or holds for more evidence; missing checks never become a pass. I would validate the prospective merged revision, promote the tested artifact through bounded exposure and retain a compatible rollback. Reports, budgets and named owners make the process usable under delivery pressure.
Remember: Define risks → Version evidence → Compare paired cases → Decide explicitly → Release the tested artifact.
Final notes: Keep hard requirements separate from average quality. Distinguish representative estimates from targeted tests. Price both builds and human review. A green check is only useful when it applies to the release actually shipped.
Related concepts: CI/CD for AI, LLM evaluation, guardrails, multi-tenant training.