Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Model Selection Guide

By Anup Rai33 min readReviewed September 2026

Model selection is choosing a model and its operating configuration to meet a defined task, objective and constraints. In a production application, the decision includes the prompt, retrieval, tools, runtime and fallback—not only the model name. This lesson concerns deployment selection; statistical model selection during training is a related but different setting.

Reviewed September 24, 2026. The catalog is a dated shortlist using primary provider documentation. Prices are USD per million tokens at direct-provider Standard rates unless noted. Numerical interview workloads and candidate outcomes are illustrative. Use model taxonomy, capability assessment and pricing and costs for deeper reference.

Architecture / visual model
flowchart TD T["Define task, correct outcome and constraints"] --> N{"Can a simpler non-generative system meet them?"} N -->|Yes| B["Keep it as a baseline candidate"] N -->|No| G["Screen model configurations against hard gates"] B --> G G --> E["Evaluate eligible complete workflows"] E --> P["Compare acceptable quality, latency and full cost"] P --> D["Choose the simplest justified configuration"] D --> C["Canary with rollback and reassessment criteria"]
Read diagram source
flowchart TD
    T["Define task, correct outcome and constraints"] --> N{"Can a simpler non-generative system meet them?"}
    N -->|Yes| B["Keep it as a baseline candidate"]
    N -->|No| G["Screen model configurations against hard gates"]
    B --> G
    G --> E["Evaluate eligible complete workflows"]
    E --> P["Compare acceptable quality, latency and full cost"]
    P --> D["Choose the simplest justified configuration"]
    D --> C["Canary with rollback and reassessment criteria"]

Define the product decision

A model is one component of a system. A support assistant also depends on current policies, retrieval, permissions, tools, and a fallback when it cannot answer. If the wrong policy reaches the model, a more capable model may produce a more persuasive wrong answer. Selection therefore starts with the task and the failure, not with a leaderboard winner.

Write three example requests before comparing candidates: a straightforward policy lookup, a difficult question requiring several facts, and a consequential action such as changing a booking. Explain what a correct result looks like for each. A model may be adequate for lookup but unreliable at selecting authorized tool arguments. The capability requirement is attached to a workflow, not simply to the word “support.”

Separate the eligibility screen from the comparison. Eligibility includes data handling, permitted regions, license terms, available modalities, and capacity. Comparison includes measured quality, latency, cost, and maintainability among eligible options. An ineligible model does not become acceptable because it wins a weighted average.

Compare eligible configurations

Suppose candidate A costs $0.02 per task and resolves 92 of 100 representative tasks, while candidate B costs $0.01 and resolves 88. These invented figures are not enough to choose. Inspect the failures. If A makes unauthorized changes while B safely hands off, their averages conceal a crucial difference. If B's additional handoffs require expensive human time, its lower model price may not reduce total cost.

Next compare latency at expected concurrency, not just one developer's sequential calls. Include long requests, output length, rate-limit behavior, and retries. A provider can accept the model ID yet lack sufficient quota for your launch. “Supports a million tokens” likewise says what input may be accepted, not whether the model reliably uses every fact or whether that request meets the response deadline.

Now decide whether routing is justified. Routing selects where a request or step executes. Pre-inference routing uses information available before generating the answer. A cascade tries one model and escalates based on a validated result check; it is one way to route a workflow. Both add components that can make mistakes. Model-generated confidence is not automatically a safe escalation signal. A schema validator, verified task result, or calibrated risk classifier may provide better evidence, depending on the task.

Finally document the choice, runner-up, assumptions, and reassessment trigger. Keep prompts, tool schemas, and evaluation cases versioned together. An abstraction layer can standardize logging and errors, but it does not make reasoning controls, caching, tool behavior, or data contracts identical across providers. Those differences need adapter tests and workload evaluation.

Decision sequence

Memorize the decision sequence: task, gates, evidence, economics, rollout. Look up the current model IDs, rate cards, quotas, and contractual terms when making the decision. The detailed catalog below is dated reference material; it is not a permanent ranking or a set of promises about your workload.

Table of Contents


The Core Principle

There is no universally best model. There is only the best validated choice for a specific workload, traffic shape, risk level, and operating environment.

Provider descriptions and public benchmarks are useful for building a shortlist. They are not a substitute for testing the exact prompt, tools, documents, languages, output schema, and failure modes used by the product.

When the product objective is to minimize cost subject to fixed requirements, use this constrained objective:

Choose the lowest-total-cost candidate that satisfies the quality, safety, latency, capacity, and governance requirements. \text{Choose the lowest-total-cost candidate that satisfies the quality, safety, latency, capacity, and governance requirements.}

The strongest model may be appropriate for a difficult agent step and wasteful for classification. The cheapest model may have the lowest token bill but the highest cost per successful task because it needs retries or human correction.


Selection Framework

Step 1: Define hard gates

Eliminate candidates that cannot meet non-negotiable requirements:

Gate Questions to answer
Modalities Are text, image, audio, video, or PDF inputs required? Is media generation required, or only understanding?
Context and output What are the P50, P95, and maximum prompt sizes? How much output can a single step require?
Tools and API features Are function calling, strict structured output, web search, code execution, computer use, MCP, caching, or batch processing required?
Deployment Must inference run in a particular region, cloud, VPC, or on-premises environment?
Data governance What are the retention, training-use, encryption, audit, and contractual requirements?
Lifecycle Is a preview endpoint acceptable? Can the application tolerate a rolling alias, or does it require a pinned version?
Capacity Can the provider sustain the required requests per minute, tokens per minute, concurrency, and burst traffic?

A large context window is only a capacity limit. It does not prove accurate recall, reasoning, or citation at that length. Test long-context behavior at the actual document mix and token distribution.

Step 2: Build candidates by operating lane

Choose a small shortlist from the lanes relevant to the task, including the existing or simplest viable baseline. A basic classifier does not need a candidate from every lane:

Lane Purpose Representative current candidates
Capability ceiling Hardest reasoning, planning, coding, and long-running agent tasks GPT-6 Astra, Claude Fable 5.1; compare GPT-6 Sol and Claude Opus 5.5 as lower-price alternatives
Balanced production Strong quality with lower cost and latency GPT-6 Sol, Claude Sonnet 5, Gemini 3.8 Flash, Grok 4.7, Mistral Medium 3.5
High-volume economy Extraction, routing, classification, simple transformation, and subagent work GPT-6 Luna, Gemini 3.5 Flash-Lite, Mistral Small 4, DeepSeek V4.1 Flash
Open-weight / controlled deployment Self-hosting, weight access, custom infrastructure, or stricter placement control Mistral Large 3, Mistral Medium 3.5, Mistral Small 4
Specialist Realtime voice, transcription, image/video generation, OCR, embeddings, or moderation Use a purpose-built endpoint instead of forcing the task through a general text model

These are candidates, not cross-provider rankings. For example, OpenAI calls GPT-6 Astra its latest flagship, while Anthropic offers Claude Fable 5.1 for demanding reasoning and long-running agents. Only a workload-specific evaluation can resolve that choice.

Step 3: Score the finalists

Use pass/fail gates before weighted scoring. A model that violates a residency requirement or fails a safety threshold cannot compensate with a better coding score.

Dimension Example measurement
Task quality Exact match, rubric score, groundedness, pass@1, tool-task completion, or human preference
Reliability Valid-schema rate, tool-call correctness, retry rate, hallucination rate, and completion rate
Safety Policy-specific false positive/negative rates, jailbreak resistance, and unsafe-action rate
Latency Time to first token, tokens per second, and end-to-end P50/P95/P99
Capacity Sustained concurrency, throttling rate, batch completion time, and quota headroom
Cost Cost per successful task, including reasoning, tools, retries, caching, and human review
Operations Version stability, observability, support, deprecation policy, regional availability, and SDK quality

Pareto dominance means one eligible option is no worse on every chosen objective and better on at least one. A Pareto frontier contains options not dominated this way. Avoid inventing one global score too early: weighted scores depend on units, normalization and stakeholder preferences. Keep hard gates visible and show the trade-off frontier: the fastest passing model, the cheapest passing model, and the highest-quality model.


Current Model Landscape

General-purpose API shortlist

The table gives shortlist candidates, not recommended winners. “Context” reports each provider’s published window/input convention; it is not a universal maximum-prompt measure. Input/output prices omit cache, tools, service, regional and contractual modifiers. Current pricing detail.

Provider Model and API ID Published role Context Input / output price
OpenAI GPT-6 Astra — gpt-6-astra Difficult reasoning and tool workflows 1.05M total; 922K max input $10 / $50
OpenAI GPT-6 Sol — gpt-6-sol Complex coding and agentic workflows 1.05M total; 922K max input $2 / $10
OpenAI GPT-6 Luna — gpt-6-luna Focused, high-volume work 1.05M total; 922K max input $0.10 / $0.50
Anthropic Claude Fable 5.1 — claude-fable-5-1 Demanding reasoning and long-running agents 1M $10 / $50
Anthropic Claude Opus 5.5 — claude-opus-5-5 Agentic coding and knowledge work 1M $4 / $20
Anthropic Claude Sonnet 5 — claude-sonnet-5 Speed/intelligence balance 1M $2 / $10
Anthropic Claude Haiku 4.5 — claude-haiku-4-5-20251001 Fastest current Claude tier 200K $1 / $5
Google Gemini 3.8 Flash — gemini-3.8-flash Current Flash model for agentic workflows and multimodal reasoning 1,048,576 input $0.75 / $3.75 through Dec. 31, 2026
Google Gemini 3.7 Flash — gemini-3.7-flash Previous-generation stable Flash; keep as an evaluated baseline, not the latest release 1,048,576 input $0.75 / $3.75 through Dec. 31, 2026
Google Gemini 3.5 Flash — gemini-3.5-flash Stable, fast multimodal and agentic model 1,048,576 input $1.50 / $9
Google Gemini 3.5 Flash-Lite — gemini-3.5-flash-lite Stable high-throughput, low-cost multimodal model 1,048,576 input $0.30 / $2.50
xAI Grok 4.7 — grok-4.7 Current flagship for code, general work, and tool use 500K $2 / $6
Mistral Mistral Medium 3.5 — mistral-medium-3-5 GA multimodal model for agentic and coding work; open weights 256K $1.50 / $7.50
Mistral Mistral Large 3 — mistral-large-2512 GA general-purpose multimodal model; Apache 2.0 weights 256K $0.50 / $1.50
Mistral Mistral Small 4 — mistral-small-2603 GA hybrid instruct/reasoning/coding model; Apache 2.0 weights 256K $0.15 / $0.60

Important constraints: GPT-6 models reserve distinct maximum input/output limits; all three allow at most 128K output. Prompts above 272K use higher whole-request rates. Claude Opus 5.5 cache reads are 5% of input, Fable 5.1 2.5%, Sonnet/Haiku 10%. Gemini 3.8 Flash’s announced January 2027 price is higher. Grok’s higher rate starts at 200K prompt tokens. These are reasons to keep the detailed, versioned rate card beside the decision.

DeepSeek is also eligible for shortlisting where its contract fits: deepseek-flash currently serves V4.1 Flash, with peak USD .30 input miss / .006 cache hit / 1.20 output; deepseek-v4-pro serves V4 Pro 0813 at 1.32 / .044 / 3.96. Off-peak rates are half. Both publish a 1M window and 384K maximum output; Flash supports image input while Pro does not. Legacy Flash aliases route to V4.1 Flash. DeepSeek model and pricing contract.

Preview and limited-access models

Keep these out of a default production baseline unless the release risk is explicitly accepted:

Model Status and consequence Published price
Gemini 3.1 Pro — gemini-3.1-pro-preview Preview endpoint with 1,048,576-token input limit. Strong candidate for multimodal and agentic evaluation, but preview lifecycle and behavior may change. $2 / $12 for prompts up to 200K; $4 / $18 above 200K
Claude Mythos 5.1 Anthropic lists it as limited availability. Do not build a general deployment plan around access that has not been contractually confirmed. $10 / $50

A provider offering an endpoint does not establish that your account has approved access, quota or contractual permission. Confirm these before planning production traffic.

Migration checks for the newest models

Candidate What can break even if the name change is simple?
GPT-6 Astra Reasoning supports low/medium/high/xhigh/max; not none/minimal. Use Responses for tools and verify supported sampling controls. EU data residency requires Standard processing.
GPT-6 Sol / Luna Reasoning also supports none. Chat Completions function calling requires none; use Responses for tool workflows with reasoning. EU data residency requires Standard processing.
Claude Opus 5.5 / Fable 5.1 Adaptive thinking is always on; forced tool selection errors. Thinking blocks are bound to model/conversation context. Validate state conversion and progress-display behavior.
Gemini 3.8 Flash Minimal thinking errors; use low/medium/high. It accepts audio/video inputs but produces text; Live/audio generation are separate models. Computer use remains a preview capability.
Grok 4.7 Text/image input and text output; low/medium/high/xhigh reasoning. Batch is not supported on this model. A generic provider-wide Batch flag would be wrong.
DeepSeek Distinguish Flash image support from Pro. Alias acceptance does not pin the retired Flash model. Test the exact endpoint and tool/schema contracts.

Sources: GPT-6 Sol, GPT-6 Luna, Astra guide, Opus 5.5, Gemini 3.8 Flash, Grok 4.7.

For new video work, do not choose Sora 2: the Videos API and Sora 2 models shut down September 24, 2026. OpenAI self-serve fine-tuning is also winding down; do not assume a new organization can train hosted custom models. OpenAI lifecycle notices.

What the table does not prove

It does not prove which model is best at coding, science, legal reasoning, long-context recall, safety, or agent autonomy. Provider-written descriptions are first-party positioning. Benchmark scores also depend on prompting, tool access, reasoning settings, sampling, grading, contamination, and harness implementation.

Use public results to select candidates. Use representative, versioned evaluations with checked grading and stated uncertainty to select a production configuration; private tests can also be biased.


Use Case Mapping

Use case Good starting shortlist What must be evaluated
Hard reasoning or professional analysis GPT-6 Astra, Claude Fable 5.1; compare Sol/Opus 5.5 for lower cost Correctness, calibration, citation quality, domain-specific failure rate, and cost of reasoning tokens
Coding agents GPT-6 Astra, GPT-6 Sol, Claude Fable 5.1/Opus 5.5/Sonnet 5, Gemini 3.8 Flash, Grok 4.7, Mistral Medium 3.5 Repository-level task completion, test pass rate, tool recovery, diff quality, security, and wall-clock time
General product assistant GPT-6 Sol, Claude Sonnet 5, Gemini 3.8 Flash, Grok 4.7 Instruction following, tone, groundedness, streaming latency, safety, and multilingual quality
High-volume extraction or classification GPT-6 Luna, Gemini 3.5 Flash-Lite, Mistral Small 4, DeepSeek V4.1 Flash Schema-valid rate, precision/recall, tail latency, batch support, and cost per accepted record
Long-document and multimodal analysis Gemini 3.8 Flash, GPT-6 Astra/Sol/Luna, Claude Fable 5.1/Opus 5.5/Sonnet 5, DeepSeek V4.1 Flash Recall by depth and position, PDF/image handling, citations, token tier changes, and truncation behavior
Realtime voice Provider-specific realtime or voice APIs End-to-end audio latency, interruption handling, transcription accuracy, voice quality, transport, and per-minute cost
Image, video, OCR, or transcription Purpose-built media/OCR/transcription models Media quality metric, resolution/duration pricing, safety filters, turnaround time, and rights requirements
Private or controlled deployment Mistral Large 3, Medium 3.5, or Small 4 weights License, hardware fit, quantization loss, serving throughput, patch cadence, and full operational TCO

Two common mistakes:

  1. Using context size as the only RAG criterion. Retrieval can reduce latency, cost, distraction, and access-control risk even when the entire corpus technically fits in context.
  2. Using a general model for a specialist workload. A text model that can reason about an audio transcript is not necessarily the right transcription engine; a vision-language model is not necessarily the right OCR system.

Evaluation and Rollout

Build an evaluation set from real traffic

Use representative examples plus a separately reported adversarial/risk suite. Targeted cases do not automatically have their production frequency. A useful initial set contains:

  • common requests weighted by real frequency;
  • rare but high-impact requests;
  • multilingual and formatting edge cases;
  • long inputs near the actual P95 and maximum;
  • malformed inputs and prompt-injection attempts;
  • every tool and output schema;
  • tasks that the current system gets wrong;
  • abstention cases where the correct behavior is to refuse or escalate.

Keep the evaluation set versioned. Separate a development set used for prompt tuning from a held-out set used for model selection.

Evaluate the complete system

For an agent, the unit under test is the full trajectory:

request→model decision→tool calls→state changes→final result \text{request} \rightarrow \text{model decision} \rightarrow \text{tool calls} \rightarrow \text{state changes} \rightarrow \text{final result}

A model can have excellent single-turn benchmark results and still fail because it chooses the wrong tool, repeats an irreversible action, loses state, or cannot recover from a tool error.

Record at least:

  • final-task success and grader rationale;
  • every model, prompt version, and reasoning setting;
  • input, cached-input, reasoning, and output token counts;
  • tool count, tool errors, and retry count;
  • time to first token and end-to-end latency;
  • provider/model error code and fallback behavior;
  • estimated and invoiced cost.

Roll out safely

  1. Offline replay: compare finalists on the held-out evaluation set.
  2. Shadow traffic: send production-shaped requests without exposing the new output or allowing side effects.
  3. Canary: expose a small, monitored cohort with explicit rollback criteria.
  4. Ramp: increase traffic only while quality, safety, latency, and cost stay within limits.
  5. Continuous evaluation: rerun the suite on every model, prompt, tool, or routing-policy change.

Pin a dated model version when the provider offers one and reproducibility is important. If only a rolling ID is available, treat provider-side changes as unannounced migrations: monitor the model identifier returned by the API and keep a rollback path. The returned name may remain unchanged even when behavior changes, so watch task-level regressions as well as identifiers.


Cost Analysis

Model the whole request

Sum cost over every execution attempt, with provider-specific non-overlapping input, cache-write, cache-read and billed-output quantities. Add distinct tool/media charges once and storage over its billing interval. Include failed work, fallback, shadow and evaluation traffic, then application operations and human review. A monthly shared storage bill is not a new fee on each request. See the executable cost calculation.

The decision metric should usually be:

Effective cost per success=total system costsuccessfully completed tasks \text{Effective cost per success} = \frac{\text{total system cost}}{\text{successfully completed tasks}}

Worked example

Assume 1 million requests per month, each with 1,000 uncached input tokens and 500 total billed output tokens, including any billed reasoning tokens. That is 1,000 million input tokens and 500 million output tokens. Actual reasoning runs may use much more output. This illustration uses current standard rates and excludes caching, tools, retries, long-context tiers, and discounts:

Model Input cost Output cost Illustrated monthly token cost
GPT-6 Astra $10,000 $25,000 $35,000
GPT-6 Sol $2,000 $5,000 $7,000
Claude Opus 5.5 $4,000 $10,000 $14,000
Claude Sonnet 5 $2,000 $5,000 $7,000
Gemini 3.8 Flash, current 2026 price $750 $1,875 $2,625
Grok 4.7 $2,000 $3,000 $5,000
GPT-6 Luna $100 $250 $350
Mistral Small 4 API $150 $300 $450

This table is not a recommendation. If the USD 450 model completes half of one million tasks, its token-only cost is USD .0009 per success. At 95% completion, the USD 2,625 model costs about USD .002763 per success. The first is still cheaper on that narrow metric, but it fails a 95% completion requirement. Quantify the additional repair/fallback costs before comparing acceptable complete systems.

Astra's standard input/output rates are 5× GPT-6 Sol's. That does not prove 5× the cost per completed job: token consumption, retries, and success rate can change. Measure them rather than assuming the newer model is either cheaper or automatically worth the premium.

Cost controls that preserve quality

  • Put stable instructions and reusable documents first to improve prefix-cache reuse; measure actual cache hits.
  • Cap output length and reasoning effort by task rather than using one global maximum.
  • Use batch or deferred processing for work without interactive latency needs.
  • Route only when routing accuracy and added latency have been measured.
  • Retrieve the smallest relevant context instead of sending an entire corpus.
  • Detect loops with per-request token, tool-call, wall-clock, and dollar budgets.
  • Track cost by tenant, feature, task type, route, and success outcome.

Operational Considerations

Rate limits are account-specific

Do not copy a static RPM/TPM table into an architecture document. Limits vary by model, usage tier, account, region, and processing mode, and they change. Read the active limits from the provider console or API, load-test below the approved quota, and maintain headroom for retries and bursts.

Plan separately for:

  • requests per minute and tokens per minute;
  • concurrent requests or sessions;
  • batch queue limits;
  • long-running streams and realtime connections;
  • tool-specific and regional quotas;
  • quota propagation time after an increase.

A common interface is not identical behavior

An abstraction layer is useful for routing and observability, but it should not pretend that providers have the same features. Maintain a capability contract for each exact deployment configuration:

Contract field What it must preserve
Identity Provider, model/version, endpoint and adapter revision
Input/output Allowed modalities, maximum input/output, schema subset and truncation semantics
Tools Tool selection, argument schema, IDs, result format and allowed feature combinations
Reasoning and sampling Supported controls, defaults and incompatible parameters
Service Region, service tier, quotas, streaming and cancellation behavior
Governance Retention, data use, approved tenants, terms and review expiry
Evidence Source reference, contract tests and evaluation run supporting approval

Use supported, unsupported and unverified states rather than assuming every provider has every feature. An unverified required feature blocks that route until checked. A basic JSON-output feature does not establish strict schema enforcement, and two individually supported features may be incompatible together.

The executable example below demonstrates request rejection. Provider adapters additionally need contract tests for roles, tool-call representation, streaming events, usage fields, image formats, errors, timeouts and cancellation.

Reliability and fallback

Use bounded exponential backoff with jitter for eligible transient throttling and 5xx errors. First classify whether a request may already have executed; do not blindly repeat an action after an unknown outcome. Respect retry headers and overall deadlines. Do not retry invalid requests, safety refusals, or non-idempotent side effects blindly.

Cross-provider fallback is a new inference path, not merely a second URL. It may need:

  • a provider-specific prompt and tool schema;
  • conversion of conversation and tool state;
  • revalidation of structured output;
  • a separate safety policy;
  • confirmation that data may be sent to the fallback region/provider;
  • a remaining latency and dollar budget.

Use circuit breakers and failure-class-aware routing. Record when a fallback occurs so degraded service does not silently become the normal service.

Governance and lifecycle

Before production approval, document:

  • data retention and training-use terms for the exact paid/free tier;
  • deployment region and subprocessors;
  • model ID, alias behavior, deprecation notice, and rollback target;
  • safety controls and human escalation path;
  • output ownership and media-rights requirements;
  • audit logging, access control, and incident response;
  • whether provider search, files, containers, or other tools retain data under different rules from the base model call.

Multi-Model Strategies

1. Evaluated routing

Route by requirements that are observable before inference: modality, context size, latency class, tenant policy, and task type. A learned complexity router can be added only after its routing errors are measured.

2. Cascade

Try a lower-cost model first, accept only when a calibrated validator says the result meets the task contract, and escalate otherwise. The validator must be tested for false acceptance; “the model sounds confident” is not a gate.

3. Specialist routing

Use dedicated transcription, OCR, embedding, moderation, realtime, image, or video endpoints for the portions they are designed to solve. A workflow can use several specialists and one general reasoning model.

4. Provider fallback

Fail over on a narrow set of transient failures. Keep provider-specific prompts and tools tested continuously so the backup path is not discovered to be broken during an outage.

5. Draft and verify

Use one model to draft and another model, deterministic checker, or human to verify high-impact work. Correlated errors matter: two calls to the same model can repeat an error, and different providers can also fail together on the same misleading evidence. Measure the verifier’s false-acceptance rate.

6. Shadow and canary

Use shadow traffic to compare a new model without user impact, then a canary to measure real-world outcomes. This is safer than changing a rolling alias for all traffic at once.

Multi-model systems add routing errors, extra latency, larger failure surfaces, more vendor contracts, and more observability work. Add them only when measured quality, resilience, or cost gains exceed that complexity.


API vs Open Weights

There is no universal request-volume crossover point. Self-hosting economics depend on model size, quantization, accelerator type, utilization, batching, latency target, redundancy, region, staffing, and the quality loss relative to the API candidate.

Prefer a managed API when

  • fastest time to market and immediate access to current models matter;
  • traffic is bursty or difficult to capacity-plan;
  • the team does not want to operate GPU scheduling and model serving;
  • managed tools, safety systems, and enterprise support are valuable;
  • workload data is permitted under the provider contract and deployment terms.

Prefer open weights or controlled hosting when

  • weights or inference must remain in a controlled environment;
  • the workload needs weight-level customization or a specialized serving stack;
  • traffic is predictable enough to keep accelerators highly utilized;
  • the team can own security patches, serving, evaluation, upgrades, and availability;
  • the chosen license permits the intended use.

Compare total cost of ownership

Include accelerators, idle and failover capacity, networking, storage, orchestration, observability, engineering/on-call time, security, model evaluation, upgrades, and the business cost of lower task quality. Benchmark with the intended precision, context length, batch size, and concurrency; a single throughput number from a different serving setup is not a capacity plan.

The current Mistral lineup offers a useful range for this evaluation: Large 3 and Small 4 use Apache 2.0 weights, while Medium 3.5 uses a Modified MIT license. Read the exact license and model card before deployment.


Manager interview practice

Decision example: two candidates pass ordinary quality tests, but one fails a required tool-permission case. Eliminate that configuration until fixed; a better mean score cannot compensate for a hard requirement. For the remaining candidate, test peak capacity and actual rework before committing.

Recall checks: Which assumptions would reverse your choice? What is the fallback's evidence? Which version, prompt, tool, and data configuration did you compare? What would justify a second provider or self-hosting?

Assign an owner to reassess on model retirement, a price change, a new region or data requirement, or a material workload shift. Memorize the method; look up the current catalog.

Executable adapter contract example

A common application interface should preserve material differences between providers. This standard-library example uses fake, already-reviewed capability contracts to show rejection before execution. Unverified contracts must not reach this constructor as approved booleans. It makes no claim about any provider's current SDK.

from dataclasses import dataclass

@dataclass(frozen=True)
class Capabilities:
    structured_output: bool
    tools: bool
    max_output_tokens: int

    def __post_init__(self):
        if type(self.structured_output) is not bool or type(self.tools) is not bool:
            raise ValueError("capability flags must be booleans")
        if type(self.max_output_tokens) is not int or self.max_output_tokens <= 0:
            raise ValueError("capability output limit must be a positive integer")

class DemoAdapter:
    def __init__(self, name, capabilities):
        if not isinstance(name, str) or not name or not isinstance(capabilities, Capabilities):
            raise ValueError("a named capability contract is required")
        self.name, self.capabilities = name, capabilities

    def generate(self, *, needs_schema, needs_tools, max_output_tokens):
        if type(needs_schema) is not bool or type(needs_tools) is not bool:
            raise ValueError("request flags must be booleans")
        if type(max_output_tokens) is not int:
            raise ValueError("output budget must be an integer")
        c = self.capabilities
        if needs_schema and not c.structured_output:
            raise ValueError("Structured output is a required capability")
        if needs_tools and not c.tools:
            raise ValueError("Tool use is a required capability")
        if not 0 < max_output_tokens <= c.max_output_tokens:
            raise ValueError("Unsupported output budget")
        return {"provider": self.name, "status": "demo_only",
                "output": None, "usage": None}

adapter = DemoAdapter("candidate-A", Capabilities(True, False, 2048))
assert adapter.generate(needs_schema=True, needs_tools=False,
                        max_output_tokens=256)["status"] == "demo_only"

In a real adapter, preserve finish reasons, truncation, refusal, tool-call IDs, cancellation behavior, and usage categories. Do not turn unsupported tool use into plain text and pretend the same contract succeeded. Run the same application acceptance cases against each concrete adapter, plus provider-specific cases for features whose semantics differ. Fallback eligibility depends on these contracts as well as availability.

Interview: select and deploy an invoice-extraction system

Problem: choose a configuration that extracts invoice fields for a reviewer. The system prepares a draft record; it does not pay invoices or change supplier bank details. All figures below are interview assumptions.

Requirements and evaluation plan

Functional requirements

  1. Accept authorized invoice uploads and extract supplier, invoice number, date, currency, subtotal, tax and total.
  2. Attach document/page evidence to each extracted field; distinguish absent, ambiguous and unreadable fields.
  3. Validate types, arithmetic and document identity outside the model, then present a reviewable draft.
  4. Escalate unsupported layouts and uncertain results while preserving the original document.
  5. Record the exact model/prompt/OCR/validator configuration and corrections for reassessment.

Non-functional requirements

  1. All document processing, tools, storage, logs and fallbacks must use the approved EU processing path.
  2. A prespecified two-sided 95% Wilson lower bound for complete-draft correctness before human review must exceed 95% on the representative holdout; evaluate critical errors separately.
  3. Automated draft-generation p95, including queueing and validation, must stay below eight seconds at a five-document/second peak for the defined one-page interactive workload. Measure human review/approval time separately.
  4. Input is bounded to 20,000 model tokens and 1,200 output tokens per interactive attempt; larger jobs use a separately evaluated asynchronous route.
  5. Authorization, field provenance and duplicate-upload handling must hold under retries and worker failures; no autonomous payment actions are permitted.

Use 1,000 independent representative held-out documents, separating related vendor templates appropriately during development/test splitting. Add a separately reported permission, misleading-content and critical-field suite. Compare complete extraction workflows with the same authoritative references, validators and budgets. A deterministic OCR-plus-rules baseline remains in the comparison.

Candidate Complete correct drafts / 1,000 Draft p95 at target load Approved EU path Interpretation
Existing OCR + rules 850 1.8 s Yes Useful baseline, fails the required quality screen
Configuration A 964 6.4 s Yes Eligible after the separate critical/operating criteria pass
Configuration B 971 10.2 s Yes Fails the interactive deadline; may suit async work
Configuration C 980 3.1 s No Ineligible regardless of quality or price

A’s 95% Wilson interval is approximately 95.06%–97.39%, narrowly passing the stated quality rule. That is evidence under the sampling assumptions, not protection against dataset bias. Do not claim A is statistically superior to B: their marginal counts alone omit paired fixes/regressions. B’s deadline failure is enough to exclude it from this interactive route. None of the scores establishes that permission leaks are impossible.

Start simple, find the failures, then improve

Initial design: upload → OCR → one model call → JSON → reviewer. It provides a fast prototype. JSON can be valid but factually wrong; uploads can be repeated; a model alias can change; a fallback can send private invoices outside the approved region; an interrupted call can leave its result unknown.

Detailed design

Architecture / visual model
flowchart TB U["Authenticated upload + document limits"] --> O[("Tenant-scoped original and digest")] O --> Q["Idempotent job + bounded queue"] R["Versioned holdout, operating tests and decision"] --> C[("Approved configuration registry")] C --> W["Worker: approved OCR + model adapter"] Q --> W W --> V["Schema, arithmetic, identity and evidence checks"] V -->|Valid draft| D[("Versioned draft record")] V -->|Uncertain or failed check| H["Reviewer / documented exception route"] D --> H H --> A["Explicitly approved record; no payment execution"] W -->|Provider outcome unknown| X["Reconcile attempt; bounded permitted retry"] X --> V D --> E["Outcome samples, corrections, latency and full cost"] A --> E E --> R
Read diagram source
flowchart TB
    U["Authenticated upload + document limits"] --> O[("Tenant-scoped original and digest")]
    O --> Q["Idempotent job + bounded queue"]
    R["Versioned holdout, operating tests and decision"] --> C[("Approved configuration registry")]
    C --> W["Worker: approved OCR + model adapter"]
    Q --> W
    W --> V["Schema, arithmetic, identity and evidence checks"]
    V -->|Valid draft| D[("Versioned draft record")]
    V -->|Uncertain or failed check| H["Reviewer / documented exception route"]
    D --> H
    H --> A["Explicitly approved record; no payment execution"]
    W -->|Provider outcome unknown| X["Reconcile attempt; bounded permitted retry"]
    X --> V
    D --> E["Outcome samples, corrections, latency and full cost"]
    A --> E
    E --> R

The registry contains the approved model, endpoint, region, prompt, OCR version, validator and fallback policy. A successful model call does not itself approve a business record. Every field links to the relevant document revision; a correction creates a new draft revision. A changed document or configuration requires revalidation.

API/state contract: POST /extractions accepts a client idempotency key under the authenticated tenant and binds it to the uploaded digest and extraction configuration. Reusing a key with different input returns a conflict. A job moves through queued → running → draft-ready or needs-review/failed. Persist attempt IDs and leases; an unknown external result is reconciled before another attempt is dispatched. A reviewer updates a specific draft revision, so a late worker cannot overwrite the correction. Idempotency in our store prevents duplicate drafts; it does not guarantee a provider executed or billed only once.

Capacity: at five documents/second and an illustrative mean active processing time of two seconds, Little’s Law gives ten active jobs on average. Provision and test bounded concurrency—for example, a candidate limit of twenty—with headroom for tails, provider quotas and one worker loss. That limit alone does not prove the eight-second p95. At the 20,000-token input bound, five calls/second could demand six million input tokens/minute before retries. Check both token and request quotas against the actual input distribution.

Failure / decision Repair and benefit Cost or limitation
Valid JSON with invented total Arithmetic and cited-source checks; inspect record correctness Evidence checks and human exceptions add work
Long or unreadable document Explicit async/review route Slower completion; separate capacity and evaluation
Configuration B is more accurate but slow Restrict it to separately approved async use, if worthwhile Additional routing and operating cost
Cheap fallback uses the wrong region Filter by the full approved processing contract before dispatch Lower availability when no eligible route remains
Alias or prompt changes Versioned release, regression cases and canary Migration and repeated assessment cost
Reviewer correction races a worker Draft revision check and append-only attempt evidence State and conflict-handling complexity
Repeated uploads spend twice Tenant-scoped idempotency and digest binding Storage/lookups; provider uncertainty remains
Router learns only from its choices Sample alternatives on permitted offline/shadow cases Extra evaluation cost; counterfactual evidence is incomplete

Full economics and decision

For 50,000 one-page documents/month, use these invented prices and operating assumptions. Exception review means additional correction work. Mandatory final approval is budgeted separately for every document at an illustrative fifteen seconds each; validate that workload too.

Monthly cost Existing OCR + rules Configuration A
OCR, USD .004/document USD 200 USD 200
Model, USD .010/document — USD 500
Exception review, two minutes at USD 60/hour 15% × 50,000 × USD 2 = USD 15,000 4% × 50,000 × USD 2 = USD 4,000
Final approval, fifteen seconds/document at USD 60/hour USD 12,500 USD 12,500
Common application operations USD 1,500 USD 1,500
Incremental adapter/monitoring operations — USD 600
Adoption/evaluation, USD 3,600 over six months — USD 600
Total USD 29,200 USD 19,900

Estimated saving is USD 9,300/month. If both processes produce 48,000 verified accepted records after review, costs are about USD 608.33 and USD 414.58 per 1,000 accepted records. Validate the four-percent review rate and accepted-outcome assumption in the canary; field accuracy alone does not establish reviewer workload.

The proposed lines excluding exception review total USD 15,900. Its TCO reaches the USD 29,200 baseline at 6,650 reviewed documents, a 13.3% review rate. Above that rate, the proposed cost advantage disappears under these assumptions. Additional proposal-specific capacity, retries or approval labor would lower that break-even rate; equal new costs on both sides cancel in this comparison. Do not transfer this result to another product without recalculating.

Rollout: approve the exact A configuration after the holdout, critical-error and load gates pass. Shadow only within the approved data path, then canary by stable tenant/document cohort. Revert on a reproduced permission violation, corrupted critical field, material deadline breach or failed agreed economic guardrail. Keep the original documents, old draft revisions and a tested reviewer path. Analyze statistical improvement at its planned horizon; a safety stop is not a claim of superiority.

Closing remarks: “A fits the interactive requirements; B is worth considering separately for asynchronous work, and C is excluded by the data contract. I would keep the initial deployment to one model configuration and a reviewer path. The estimated saving is mostly reduced exception work, so the canary must verify that assumption. I would add a router only if its measured benefit exceeds the extra errors, state handling and operating cost.”

Interview questions with developed answers

Q1: How do you choose between models from OpenAI, Anthropic, and Google for production?

Show answer and follow-up

Sample answer: I translate the product into representative tasks and hard constraints first. I shortlist eligible models, then compare complete configurations on the same held-out cases, including difficult and prohibited-action scenarios. I inspect failure severity and slices, benchmark latency and quota behavior under realistic load, and calculate cost per acceptable completed task including escalation. I choose the lowest-cost configuration that meets the requirements, document the tradeoff, and validate with a canary. The dated model catalog helps find candidates; it does not replace that evidence.

Follow-up: What if one candidate needs a different prompt? Tune each reasonably and disclose that the comparison is between configured systems.

Q2: When would you self-host rather than use an API?

Show answer and follow-up

Sample answer: I consider self-hosting when control requirements, an appropriate available model, predictable demand, and operating capability justify it. I include GPU capacity, spare capacity, storage, network, upgrades, security, and on-call staffing in total cost. An API can be better for uncertain demand or access to capabilities we cannot operate ourselves, provided data and service requirements are met. There is no universal queries-per-month crossover because token lengths, batching, utilization, and quality differ. I would benchmark the exact workload and compare equivalent service objectives before recommending a migration.

Follow-up: Does self-hosting guarantee privacy? No; access control, logs, backups, and network paths still require engineering.

Q3: When is a multi-model router worth building?

Show answer and follow-up

Sample answer: A router is useful when task groups have meaningfully different requirements and the savings exceed its engineering and error costs. I need labels that describe which candidate can complete each task acceptably, a policy for uncertain classifications, and evaluation of the complete routed system. I would begin with a few understandable task categories, compare against a single-model baseline, and track escalation and routing errors. If the workload is small or mostly homogeneous, the router can add more maintenance and latency than value.

Follow-up: How do you avoid a feedback loop? Continue sampling alternatives and independently grading outcomes instead of learning only from the model the router already chose.

Q4: A model wins a benchmark but loses your evaluation. Which do you trust?

Show answer and follow-up

Sample answer: I investigate the mismatch before choosing. The public benchmark may test a different task, use a different budget, or omit retrieval and tools. Our own set may also be biased or incorrectly graded. I compare protocols and inspect the actual failures. Once the workload set and grading are credible, product-specific performance drives the decision. I retain the benchmark as evidence about the capability it measures, not as a universal verdict that overrides our users' requirements.

Follow-up: What would make you reject your internal result? Leakage, unrepresentative cases, inconsistent tuning, or unreliable labels.

Q5: How do you prepare for provider model updates?

Show answer and follow-up

Sample answer: I record model identifiers and relevant configuration, track lifecycle notices, and keep representative regression cases and a validated fallback. When the provider supports version pinning, I use it where stability matters, while recognizing that availability and service behavior still need monitoring. I evaluate proposed changes, canary the whole workflow, and watch task outcomes and cost. I also budget recurring migration work. Portability is the ability to move with understood changes and evidence, not merely changing a string in a client library.

Follow-up: What if an alias changes unexpectedly? Contain risky actions, compare behavior against the recorded baseline, and use a tested alternate configuration where permitted.

Q6: A candidate leads quality and cost but uses an unapproved region. Can a weighted score select it?

Show answer and follow-up

No. The processing requirement is a hard gate, including tools, logs, storage and fallback. Remove the ineligible configuration before scoring preferences. If the business wants to change that requirement, the relevant owner must explicitly change it; the routing algorithm cannot silently waive it.

Follow-up: Does self-hosting automatically solve this? No. You still need approved access, network, logs, backups and operations.

Q7: Does a 1.05-million-token window accept a 1-million-token prompt?

Show answer and follow-up

Not necessarily. GPT-6 Astra, Sol and Luna publish a 1,050,000 total window but a separate maximum input of 922,000 and maximum output of 128,000. Check the model’s input cap, output cap and combined context accounting. Even an accepted prompt may miss the required latency or useful-context accuracy.

Follow-up: Can you compare all vendor context numbers directly? First distinguish total-window and input-limit conventions, then test the actual workload.

Q8: The provider supports tools and reasoning. Why does the migrated call fail?

Show answer and follow-up

Feature combinations and endpoint contracts matter. For example, GPT-6 Sol/Luna function calling through Chat Completions requires reasoning set to none; tool workflows with reasoning use Responses. Preserve supported controls, schema subsets, service tier and region in the adapter contract. A provider-wide supports-tools flag is insufficient.

Follow-up: Should an adapter quietly remove the unsupported option? Only when the application explicitly permits the changed semantics; otherwise reject the configuration.

Q9: Configuration A has 964 correct drafts and B has 971. Which is better?

Show answer and follow-up

The point estimates alone do not establish a quality difference. Inspect paired outcomes, severity and the planned uncertainty analysis. In the example B also misses the interactive p95 deadline, so it is excluded from that route even though its observed accuracy is higher. It could be tested separately for asynchronous use.

Follow-up: Why does A pass the illustrative quality criterion? Its two-sided 95% Wilson lower bound is about 95.06%, just above the prespecified 95% requirement under the sampling assumptions.

Q10: Can a model grader make a cheap-first cascade reliable?

Show answer and follow-up

Only if its acceptance decisions are validated for the task. Measure false acceptance, escalation and correlated failures. A fluent wrong invoice total can fool both generator and judge. Prefer deterministic arithmetic and document-evidence checks where possible, with human review for unresolved cases. Include validator and fallback costs.

Follow-up: Is a different provider an independent verifier? Not automatically. Shared training patterns, prompts or wrong evidence can correlate their errors.

Q11: Why retain a deterministic baseline during model selection?

Show answer and follow-up

It can already satisfy simple tasks with predictable cost, latency and behavior. It also reveals whether adding a model improves the actual outcome. Compare the full baseline workflow, including exceptions and human correction, against the proposed system. A model is justified by a measured benefit, not merely by the product being called AI.

Follow-up: Can the baseline fail the automation target yet remain operationally useful? Yes, if a staffed exception path completes the process; account for that extra work separately.

Q12: A fallback is healthy but has no validated tool-state conversion. Should you use it during an outage?

Show answer and follow-up

Not for a workflow requiring that contract. Reject unsupported state, reconcile any uncertain external action and use a permitted degraded or human path. Test fallbacks before incidents, including tool-call IDs, refusal, truncation, usage and region. A second endpoint is not a complete continuity plan.

Follow-up: Can you rerun a timed-out payment step with a new ID? No. Reconcile the original operation before any approved retry; a new model cannot determine the actual ledger outcome from a timeout alone.

Q13: The model alias has not changed. Is another regression run unnecessary?

Show answer and follow-up

No. Returned identifiers can stay the same despite serving or behavior changes. Keep representative outcome monitoring, controlled comparisons and lifecycle notices. Pin versions where available, but also monitor prompts, retrieval, tools, defaults and operating conditions. Reassess on material changes rather than on name changes alone.

Follow-up: Does version pinning reproduce every answer exactly? No. It does not by itself guarantee deterministic sampling, identical hardware execution or unchanged external tools.

Q14: The token bill increases by USD 500. Why might the new system still be worthwhile?

Show answer and follow-up

The invoice example reduces exception work by USD 11,000/month while adding USD 500 model cost and USD 1,200 incremental operations/adoption cost. Total falls from USD 29,200 to USD 19,900 under the stated assumptions. Validate review rates and accepted outcomes; do not assume accurate-looking JSON saves labor.

Follow-up: Where does the estimated cost advantage disappear? At a 13.3% exception-review rate, holding the other scenario assumptions fixed.

Q15: When should a team add a second production model?

Show answer and follow-up

When measured quality, modality, resilience or cost benefits justify the extra routing, contracts, evaluation and incident burden. Start with clear task classes and a single-model baseline. Include wrong-route errors and sample alternatives so the router does not learn only from its own choices. For the worked interactive use case, one approved configuration plus a reviewer path is a defensible starting point.

Follow-up: What belongs in the closing recommendation? Chosen configuration, excluded options and reasons, full economics, unresolved assumptions, canary metrics, rollback owner and reassessment trigger.

Final recall table and notes

Step What to say in an interview
Task Define a correct, verified outcome and the simplest baseline
Gates Eliminate incompatible data, feature, capacity and lifecycle contracts
Evidence Compare configured workflows on held-out tasks; inspect uncertainty and severity
Economics Include every attempt, review, operation and migration cost
Design Preserve identity, state, permitted fallback and provider differences
Rollout Canary the exact release; name rollback criteria and a reassessment owner

Tip: make the decision explicit. “A is the current choice because it passes these requirements at this total cost; B is excluded by latency; C is excluded by data policy” is more useful than ending with a list of vendor names. State assumptions that could change the answer. A new model release is a reason to evaluate, not automatic evidence to migrate.

60-second interview answer

I choose a production configuration by hard requirements first, then workload evidence. I shortlist models that meet data, modality, lifecycle, and capacity constraints, compare quality and serious failures on the same representative tasks, and measure complete-task latency and cost. I choose the simplest passing option, document the tradeoff, and canary it with rollback. A public benchmark or provider description helps identify candidates; it does not establish the best choice for our application.

Remember: Gates → workload tests → total cost → controlled rollout.

Official Sources

Model contracts and rate cards checked September 24, 2026:

Next: Fine-Tuning Guide

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Pricing and Costs
NEXT LESSONPretraining: learning a reusable language model →

Explore the diagram