An ensemble combines information from multiple predictors or model outputs to produce a decision. In LLM applications, this can mean voting over answers, selecting a candidate using scores, combining drafts or aggregating judgments. It is a technique to evaluate, not a requirement for every production system.
The potential benefit comes from useful complementary information. If every model repeats the same wrong assumption, adding more votes does not repair it. Different model names do not establish independent errors.
Distinguish the common patterns
| Pattern | What runs | How the result is chosen | Typical failure |
|---|---|---|---|
| Self-consistency | Several sampled solutions, often from one model | Aggregate equivalent final answers | Shared misconception wins the vote |
| Best-of-N | N candidates plus a verifier/scorer | Select the best valid candidate | Scorer prefers a convincing wrong answer |
| Judge panel | Several evaluators grade the same answer | Aggregate compatible judgments | Judges share bias or misunderstand the rubric |
| Debate/critique | Models inspect and revise answers | A selection or synthesis rule resolves outputs | Models converge on persuasion instead of evidence |
| Mixture of agents | Proposers feed outputs to one or more aggregators | Generate a synthesis from prior outputs | Synthesis introduces new unsupported claims |
| Routing | A policy chooses an eligible model | Usually one initial generator runs | Router misclassifies the task |
Routing and fallback are related orchestration choices, but do not necessarily combine multiple predictions. See gateway routing. “Arbitration” describes choosing among candidates; it is not a universally separate category that forbids selection within an ensemble design.
A mixture of agents coordinates model calls at application level. A mixture-of-experts model routes computation among expert components inside a model. They are different architectures.
Interview scope: generate and review coding exercises
A learning service produces a solution and explanation for a programming exercise. Hidden tests and a human rubric establish whether it is useful. Model-written tests alone are insufficient: the same misunderstanding can influence both solution and tests.
Functional requirements
- Generate candidates within a common task specification and allowed context.
- Reject invalid or disallowed candidates before quality selection.
- Run isolated executable checks where possible and evaluate explanations separately.
- Return one selected answer, or abstain/escalate when no candidate qualifies.
- Record candidate, scorer and selection versions for later comparison.
Non-functional requirements
- Bound total generation, judging, retries, latency and spending per task.
- Keep all providers within the task's access and data-processing constraints.
- Prevent candidates from executing external writes while being compared.
- Measure actual task success and error correlation on held-out examples.
- Preserve a useful single-model baseline for quality and cost comparison.
Start with one generator and deterministic tests. Add more candidates only when the baseline's error analysis suggests a benefit that a stronger single model, better retrieval or a better prompt does not provide more efficiently.
Self-consistency: aggregate answers, not confidence claims
The original self-consistency research samples multiple solution paths and aggregates the resulting answers. For an arithmetic task, normalize final values and units before counting: 0.5 seconds and 500 ms should not become opposing votes. A unitless 500 may be ambiguous and should not be normalized by guessing.
For open-ended writing, exact string voting usually makes little sense. Semantic grouping is possible but introduces a grouping model or rules that can themselves make errors. Do not impose a fixed temperature or sample count as a universal recipe.
Worked probability: assume each of three binary classifiers is independently correct with probability 0.7 on a case. Majority correctness is:
Under the same assumptions, five classifiers reach 0.83692. These are calculations under an idealized model, not forecasts for a real LLM panel.
| Assumption changes | Consequence |
|---|---|
| All members make exactly the same errors | Majority accuracy remains 0.7 |
| Members are systematically wrong on a topic | More votes can reinforce the wrong answer |
| Several valid outputs are treated as different labels | Voting can reject useful answers |
| A few members time out | The aggregation rule must specify whether fewer votes are acceptable |
A 4/5 vote share is agreement, not an 80% probability that the answer is true. Calibrate any confidence estimate against labeled outcomes for the actual workflow.
Best-of-N: the selector matters as much as generation
Generate N candidates, apply hard eligibility checks, then rank candidates using a verifier or calibrated rubric. For code, use controlled tests and inspect security and maintainability separately. Passing an incomplete test suite does not prove a general specification.
Best-of-N can apply to mathematics as well as prose. A strong verifier may outperform simple majority voting. Conversely, a weak preference scorer can select fluent wrong answers from a better candidate set.
Optimizing harder against an imperfect score can worsen true quality. Reward-model overoptimization research studies this behavior, including best-of-N selection. Multiple reward models or conservative score aggregation can help in a measured setting, but neither guarantees immunity to reward hacking. Keep independent outcome checks and held-out tests.
Judge panels and pairwise evaluation
A panel grades the same output using a defined rubric. Normalize scales and handle missing/invalid judgments; do not silently average unrelated score meanings. Use median or trimmed means only when appropriate for the scale and sample count. Trimming the highest and lowest of two scores leaves no data.
The PoLL paper reports benefits from a diverse panel in its evaluated settings. It does not establish that every panel of small models beats every larger judge or that family labels alone measure diversity.
For a pairwise comparison, randomize answer order and, where the evaluation budget warrants it, compare both orders. Map choices back to candidate IDs before aggregation. If judgments disagree, the cause may be position bias, stochastic variation, ambiguity or a genuine tie; it is not automatically proven positional bias.
Do not calculate 1 - standard_deviation / mean and label it confidence. It can be negative, depends on score scale and has no general probabilistic interpretation. Track agreement separately from agreement with expert labels.
Debate and synthesis need external evidence
Critique can expose missing assumptions and tests. It can also propagate one model's error through the group. Keep initial answers independent before sharing them, require evidence or executable checks for disputed claims, and cap rounds by the task budget. There is no universal optimal two-round debate.
Mixture-of-agents research uses layered outputs as auxiliary information for later models. A production synthesis must preserve source provenance and validate the final answer; validating only the input drafts is insufficient. Compare that design against the simpler option of selecting an already valid candidate.
Evolve the baseline into a bounded ensemble
Read diagram source
flowchart TD
Q[Task, authorized evidence and shared budget] --> A[Independent candidate A]
Q --> B[Independent candidate B]
Q --> C[Independent candidate C]
A --> V[Schema, policy and isolated task tests]
B --> V
C --> V
V --> E[Score eligible candidates with versioned rubric]
E --> S{Any candidate meets acceptance rule}
S -->|No| H[Abstain, clarify or review]
S -->|Yes| R[Select answer and validate final presentation]
R --> U[Return one result]
If the chosen result proposes a real action, execution occurs later through a single authorized action path with a stable operation ID. Running three candidates must not create three tickets, payments or publications. Retries and cancellation must account for work already running or already billed.
| New problem | Repair | Cost of the repair |
|---|---|---|
| Same retrieval omission affects all candidates | Improve evidence and test source diversity | Additional retrieval and source validation |
| Judges prefer verbosity | Calibrate rubric and inspect length effects | Expert labels and maintenance |
| One provider times out | Define minimum evidence or a valid fallback | Reduced quality/coverage or greater latency |
| Candidate text manipulates the judge | Treat candidate as untrusted; isolate instructions and use independent tests | Detector/verifier work; residual risk remains |
| Majority answer violates a hard rule | Apply the rule before selection | Possible abstention despite agreement |
| Synthesis changes a correct number | Validate the final synthesis against evidence | Another check on the critical path |
Cost and latency: calculate the whole path
Illustrative prices per call:
- Four generators at $0.004 each: $0.016.
- Four candidate checks at $0.0005 each: $0.002.
- Final selection/presentation step at $0.001: $0.001.
- Total: $0.019 per task, or $950 for 50,000 tasks, before retries, infrastructure and review.
A single $0.004 generator costs $200 for the same volume before its checks. Compare cost per accepted outcome with the same acceptance standard; a higher raw quality score alone does not establish better economics. Apply the ensemble selectively if a measured routing rule identifies tasks that benefit.
Parallel execution reduces serial waiting, but waiting for all candidates still follows the slowest response. Approximately:
If five independent calls each finish within three seconds with probability 0.95, all five do so with probability . Shared congestion makes independence questionable and can worsen behavior. Fan-out also consumes more quota and concurrency, so parallelism does not guarantee single-call latency.
Run the comparison fairly
- Freeze a representative held-out task set with important language, difficulty and risk slices.
- Compare single model, stronger single model, repeated sampling, selected panel and any synthesis variant.
- Keep evidence access, allowed tools and task budgets explicit.
- Measure accepted outcomes, severe errors, abstention, human review, p95 latency and total cost.
- Inspect paired wins/losses and uncertainty rather than reporting only average score differences.
- Ship only the complexity supported by measured benefit; continue checking after model or traffic changes.
Interview tip: Explain which errors your extra candidate is expected to correct. “Three agents are more reliable” is not a mechanism.
Interview questions and answer checks
| Question | A strong answer includes |
|---|---|
| When does majority voting help? | Useful diversity, suitable answer aggregation and sufficient individual quality; explicit independence limitations |
| Why is unanimous agreement not proof? | Shared training, evidence and prompts can produce correlated errors |
| Can best-of-N help on math? | Yes, when a verifier ranks valid solutions better than a vote; benchmark it |
| What if all candidates fail the schema? | A bounded repair or fallback; do not select the least invalid candidate as valid |
| Why can more samples reduce true quality? | Selection overoptimizes an imperfect score and finds its weaknesses |
| Does swapping answer order remove all judge bias? | It diagnoses some order sensitivity; calibration and external labels are still needed |
| Is a judge panel a security boundary? | No; enforce permissions and hard constraints independently |
| How do partial timeouts affect a panel? | Predefine required evidence, deadline, missing-vote handling and acceptable degradation |
| Why keep writes out of candidate generation? | Comparison should not multiply real-world effects; execute the selected authorized operation separately |
| When should the design return to one model? | When additional cost/latency/complexity lacks a measured outcome benefit |
Final notes
Remember diversity → validation → aggregation → outcome measurement. An ensemble is useful when it corrects relevant errors at an acceptable cost. State the assumptions behind voting, validate the selector and final answer, and keep permissions and external effects outside the competition among candidates.