A coding model predicts useful code or tool calls; a coding agent combines a model with context, tools, an execution loop, and verification. A product adds a user interface, account policies, integrations, and billing. These are different decisions, even when one vendor supplies all of them.
For an interview, explain how you would select and operate a coding system. A leaderboard position or an editor preference is not an architecture argument.
Separate the decisions
| Decision | Examples of what you choose | What it does not establish |
|---|---|---|
| Model | Hosted model, open-weight model, exact revision | Repository access, safe execution, correctness |
| Agent runtime | Tool loop, context management, retry/stop policy | An effective model for your workload |
| Workspace | Local checkout, isolated remote worker | Permission to merge or deploy |
| Interface | Editor, terminal, browser, API | Where computation or data actually goes |
| Verification | Tests, analysis, review, acceptance criteria | That every possible defect has been excluded |
An editor can launch a remote agent. A terminal agent can call a hosted model. Running the agent locally does not mean inference is local. Trace source code, prompts, tool results, telemetry, and credentials separately.
Modern options: capabilities overlap
The following is a documentation snapshot reviewed in September 2026, not a ranking or a price list. Check the actual plan, version, and deployment mode before adopting a capability.
| Option | Relevant role | Question to investigate |
|---|---|---|
| Claude Code | Coding agent across several interfaces; separate Agent SDK | What permissions and execution isolation apply? |
| Codex | Coding agent with CLI, IDE and hosted workflows, plus programmatic integration | Which environment holds the checkout and runs tools? |
| GitHub Copilot cloud agent | Repository work in an ephemeral GitHub Actions environment | Which workflows, credentials and branch rules can it reach? |
| Cursor | Editor experience, agent workflows and additional CLI/cloud surfaces | Which features and policies are included in this deployment? |
| Cline | Agent tooling across editor and terminal surfaces, with an SDK | Who supplies the model and pays inference charges? |
| OpenHands | Agent SDK/server and browser-based Agent Canvas; managed and self-hosted options | Which component, license and isolation boundary are you adopting? |
| Aider | Repository-oriented coding assistance with a token-budgeted repository map | Does the selected model work well with its edit/context workflow? |
| Google Antigravity | Agentic development surfaces including desktop, CLI and IDE integrations | What are the supported account, execution and model options? |
The Windsurf documentation entry now leads to Devin Desktop documentation; verify the current product and contract rather than relying on an old Codeium feature/price table. Do not describe competing tools as “autocomplete only” without checking their current agent capabilities.
Aider's repository map selects useful symbols and relationships within a token budget; it does not place every source file in the prompt. OpenHands offers multiple components, so a single historical Docker command or license label is not a complete description. A freely available client can still incur model, compute, storage, and operational charges.
Open weights: inspect the artifact, not just the family name
Open weights means model weights are available under specified terms. It does not automatically mean the training data is available, every family member has the same license, or a deployment satisfies all organizational requirements.
Two current examples illustrate different operating profiles:
- Qwen3-Coder-Next documents 80 billion total parameters, about 3 billion active per token, a hybrid architecture, and Apache-2.0 licensing. Its card describes a non-thinking model; do not assume every Qwen model uses the same reasoning format.
- DeepSeek-V4-Pro documents a much larger mixture-of-experts model and mixed weight precision. Its weight card labels the artifact a preview. A hosted service's release status and a particular downloadable artifact's status need not match.
These examples are candidates to evaluate, not universal recommendations. Older coding models can still be useful for a constrained workload, but “latest,” “best,” and “commercially unrestricted” require evidence for the exact artifact.
Before serving one, record:
- Model repository, revision, license and applicable usage terms.
- Required tokenizer/chat template and tool-call format.
- Supported serving engine, hardware and precision.
- Measured quality on your languages, repositories and task types.
- Context length and concurrency at the target latency.
- Upgrade, rollback, monitoring and security ownership.
Total parameters and active parameters answer different questions
Mixture of experts routes a token through part of the model. Active parameters help explain computation; they are not the total weight-storage requirement. Check the model's actual placement/offloading design.
For an idealized 80-billion-parameter model with every weight stored at four bits:
raw weight bytes = 80,000,000,000 × 4 / 8
= 40,000,000,000 bytes ≈ 40 GB
That is not a claim that the model fits or serves well on a 40 GB device. Quantization metadata, unquantized components, caches, runtime buffers, fragmentation, and concurrency add requirements. CPU offload changes bandwidth and latency. See inference and serving for the capacity tradeoffs.
Understand what benchmarks measure
| Benchmark or measure | Useful evidence | Important limitation |
|---|---|---|
| HumanEval / HumanEval+ | Function-level code generation; EvalPlus adds more tests | Does not represent an entire repository workflow |
| LiveCodeBench | Coding tasks with release dates; several code-related scenarios | A date filter helps only when the relevant model training cutoff is known |
| SWE-bench Verified | Resolving a human-filtered set of 500 repository issues | Score depends on model and agent setup, tools and budget |
| Internal task suite | Your languages, dependencies, failures and acceptance criteria | Requires maintained, representative tasks and independent evaluation |
EvalPlus strengthens correctness testing and also provides efficiency evaluation. LiveCodeBench supports time-based subsets; it does not prove that every model evaluated on every subset has never encountered the problems. The SWE-bench site distinguishes evaluation tracks, including a standardized Bash-only setup. Read the track before comparing scores.
Pass@1 estimates success with one sampled solution; pass@k estimates the probability that at least one of k candidates succeeds under the stated sampling/evaluation setup. A product must still choose which candidate to use. “Resolved percentage” on a repository benchmark is not interchangeable with function-level pass@1.
A defensible comparison records the dataset revision, model version, agent/harness version, allowed tools, token/time limits, sample count, environment, and evaluation date. Do not compare one model with ten attempts against another with one and attribute the difference entirely to model quality.
Interview tip: Report “this configuration resolved these tasks under this budget,” not “this model is 87% good at software engineering.”
Interview design: an internal coding-assistant service
Functional requirements
- Accept an authorized repository, immutable starting revision, and scoped task.
- Create an isolated workspace and inspect relevant code.
- Produce a patch, optionally run approved checks, and explain its evidence.
- Support cancellation, bounded retries, and resumable job status.
- Submit the result through the team's normal review/release workflow.
Non-functional requirements
- Isolate organizations, repositories, credentials and concurrent tasks.
- Enforce limits on execution time, model spend and worker resources.
- Preserve the starting revision and complete patch provenance.
- Measure queue time, completion time, accepted changes and regressions.
- Keep source and logs within the chosen data-handling boundaries.
Start with one queue, one worker type, one model, and a narrow task category. A fleet of specialist agents is not necessary to prove value.
Read diagram source
flowchart LR
U[Authorized task] --> A[API and repository authorization]
A --> Q[Job store and queue]
Q --> W[Isolated worker at pinned revision]
W --> L[Model endpoint]
L --> W
W --> T[Scoped tools and bounded execution]
T --> W
W --> P[Patch and evidence artifacts]
P --> V[Independent verification]
V --> R[Repository review and release policy]
W --> O[Redacted status and usage events]
Repository text is untrusted input to the agent. Policy enforcement belongs in the service and tool boundary. Do not interpolate issue bodies into shell source. Give workers short-lived, scoped access and keep deployment credentials out of ordinary coding jobs. A mounted Docker control socket can grant powerful host access; “it runs in a container” is insufficient evidence of isolation.
Find the flaws, then improve the design
| Failure | Improvement | Cost or limitation |
|---|---|---|
| Worker crashes after producing a patch | Persist artifacts and job transitions; resume from a known revision | More storage and explicit recovery states |
| Duplicate queue delivery repeats an external write | Stable operation ID, conditional job claim and idempotent submission | Must reconcile an unknown submission outcome |
| Agent edits tests until they pass | Preserve independent acceptance checks outside its write scope | Requires maintained verification fixtures |
| Concurrent jobs change the same file | Separate workspaces; revalidate against the integration revision | Rebase conflicts and extra test runs |
| Long tasks occupy every worker | Per-tenant quotas, fair queueing, timeouts and cancellation | Some tasks must wait or be split |
| A cheaper model needs many repair cycles | Route by measured task difficulty; cap escalation | More routing/evaluation complexity |
| Agent reports success without checking | Derive status from recorded checks and artifact results | Some correctness still needs expert review |
For security fixes, add targeted exploit/regression tests. For a UI task, verify the actual interaction and layout. Textual similarity to a reference patch is not a reliable correctness measure: multiple implementations can satisfy the same contract.
Cost and capacity: show the assumptions
For hosted inference, add input, output, cached-token, tool and execution charges using the provider's actual billing units. For self-hosting, include utilized and idle capacity, engineering, reliability, storage and data transfer. Neither route is automatically cheaper or compliant.
Hypothetical comparison, excluding costs shared by both routes:
hosted variable cost per accepted task = $0.30
self-hosted fixed monthly cost = $3,000
self-hosted variable accepted-task cost = $0.05
break-even accepted tasks/month = 3,000 / (0.30 - 0.05)
= 12,000
This assumes equal acceptance quality and enough self-hosted capacity. If the quality differs, compare total cost per accepted, non-regressing change, including failed attempts and review/rework. Do not use a token price as the entire cost of an engineering outcome.
For worker sizing, suppose arrival rate is 30 jobs/hour and mean worker occupancy is 10 minutes. Mean in-progress demand is 30 × 10/60 = 5 workers under stable conditions. Five workers would leave no average spare capacity; bursts and latency goals require headroom. Model requests within a job may overlap or block, so size inference capacity separately from job workers.
Interview questions and answer checks
- Why can two agents using the same model have different success rates? Context selection, tools, editing format, stopping logic, verification and budgets differ.
- Does a 3B-active MoE fit wherever a dense 3B model fits? No. Total weights, precision, placement, cache and runtime requirements still matter.
- Is a local IDE agent an on-premises inference solution? Only if its actual model and integration data paths stay there; inspect each path.
- Why might a high SWE-bench score be insufficient for adoption? Your languages, task distribution, permissions and cost limits can differ substantially.
- Can the agent run its own tests? Yes, but those results need independent acceptance checks for claims the agent could otherwise manipulate.
- Should every failed task be retried with a larger model? No. Missing permissions, broken dependencies and unclear requirements need other repairs.
- How would you evaluate a generated refactor? Behavioral equivalence where required, regression tests, performance checks when relevant, and review of the diff—not matching the reference patch's text.
- When would you self-host? When measured economics or deployment/control requirements justify the operational burden and the chosen model meets the workload's quality and latency needs.
Final notes
Remember five questions: Which model? Which loop? Which workspace? Which checks? Which cost? Pin versions and evaluate the complete configuration. Start with bounded tasks, inspect real changes, and expand autonomy when observed outcomes justify it. The strongest interview answer connects each tool choice to a requirement and explains what evidence could reverse that choice.