Learnastra AI SYSTEM DESIGNAnup Rai

Interview toolkit

Common Pitfalls in AI System Design Interviews

By Anup Rai13 min readReviewed September 2026

A weak design answer often contains correct terms but omits the reasoning that connects them. This chapter helps you identify the missing requirement, mechanism or evidence, then repair the explanation. It is a practice checklist, not a universal employer scoring rubric.

Remember: requirement → mechanism → failure → evidence → decision.

For example, a policy assistant returns an obsolete leave rule. If ingestion missed the new edition, changing embeddings cannot retrieve it. If the current edition was retrieved but its exception was removed during context packing, a larger model still lacks the qualification. If the answer exposes restricted content, factual correctness does not make the disclosure acceptable.

Architecture pitfalls

1. Skipping the data pipeline

Gap: the diagram begins with a populated database and never explains how useful information arrives.

Repair: show ingestion, parsing, content boundaries, metadata, permissions, versioned indexing, updates and deletions. Name the source owner and the behavior when a source is stale or unavailable. Preserve evidence through each transformation. See data engineering for AI.

Check: trace one corrected policy from the source to a new answer and to any previously cached answer that depends on it.

2. Choosing models by brand or adding routing automatically

Gap: one model is declared best for all tasks, or a router is added merely to mention several models.

Repair: compare eligible deployments against the workload, quality, data, latency and complete-cost requirements. A single model can be an appropriate baseline. Routing is worthwhile only when its measured benefit exceeds classification mistakes, fallback calls and operating complexity. See model selection.

Check: identify the evidence that would make you remove the router.

3. Leaving evaluation until after the design

Gap: “we will monitor accuracy” does not say what a correct result is or how failures are counted.

Repair: define the task, target population, acceptance criteria and important slices. Evaluate the baseline and proposal on comparable held-out cases; report missing outcomes and uncertainty. Monitor production behavior and business outcomes separately from offline scores. See evaluation.

Check: could a system that abstains on every difficult question improve your metric without helping users?

4. Treating a tenant field as isolation

Gap: a client-supplied tenant_id is trusted, or all candidate text reaches the model before filtering.

Repair: derive identity and scope from authenticated server context. Enforce policy before protected data reaches unauthorized consumers, including models, rerankers, logs and caches. Include current user/group permissions within a tenant, resource ownership and noisy-neighbor limits. See access control.

A database's internal candidate filtering strategy is not by itself the authorization boundary. Explain which trusted service may access which data and where enforcement occurs. Tenant-specific embedding spaces do not replace access control.

Check: use two users in the same tenant with different document access, then revoke one permission.

5. Calling any fallback graceful degradation

Gap: retrieval fails, so the system generates an unsupported answer with a disclaimer; a protected service fails, so it returns a stale cached answer.

Repair: keep the original constraints during degradation. Valid options may be authorized source passages without synthesis, a permitted alternative model, an explicit unavailable response or human escalation. Classify errors, share the deadline across attempts and prevent retry storms. See reliability patterns.

Check: which fallback is still acceptable when current authorization cannot be established?

Technical knowledge pitfalls

6. Confusing representations with evidence

Gap: an embedding is treated as verified knowledge, or similarity as factual support.

Repair: an embedding represents an input as a vector useful for learned comparisons. In a typical text RAG pipeline, the generator receives selected passages, not the search vectors as a substitute for those passages. Related passages may contradict the proposed answer. See embeddings.

Check: explain how a retrieved passage supports the particular claim, including exceptions and edition.

7. Confusing context capacity with comprehension

Gap: a context limit is treated as proof that every fact inside that limit is used correctly, or a short prompt is claimed to eliminate all position effects.

Repair: count the full request using the model/API's rules, including tools, history and output allowance where applicable. Test relevant evidence among distractors and conflicting sources. The useful context depends on the task; there is no universal safe number of chunks. See context engineering.

Check: can the answer still be recovered when its evidence appears in a different position?

8. Reporting token savings as business savings

Gap: a rate is quoted without input/output units, retries, cache writes, tool charges or review work.

Repair: calculate the same workload and accepted outcomes, with explicit rate conditions and all incremental costs. Streaming changes delivery; it is not automatically a discount. Current rates belong in the dated pricing guide.

Worked example, using hypothetical rates: 10,000 requests/day, each with 2,000 input and 500 output tokens. At $2/$10 per million input/output tokens, model cost is $0.009/request = $90/day = $2,700 per 30-day month.

If 1,500 input tokens actually hit a cache priced at $0.20/million and the remaining 500 cost $2/million, each request costs $0.0003 + $0.001 + $0.005 = $0.0063. That is $1,890/month in this idealized steady state, a $810 or 30% call-cost reduction. Cache writes, misses and provider eligibility are excluded and must be added for a real estimate.

If the cache change adds $900/month of engineering and operations, total incremental comparison is $2,700 versus at least $2,790. Even before write/miss costs, it does not save money under these assumptions. Quality and authorization must remain acceptable too.

9. Listing RAG stages without explaining choices

Gap: “chunk, embed, retrieve, rerank” is recited without saying what each stage fixes.

Decision Useful comparison Evidence
Chunking Structural boundaries versus fixed windows Evidence completeness and retrieval/answer results
Retrieval Lexical, dense or hybrid Candidate recall on relevant query types
Reranking Existing ordering versus a second-stage ranker Relevant evidence reaches the packed context
Context packing More passages versus focused evidence Answer support, latency and token cost
No evidence Limitation versus additional bounded search Coverage improvement without unsupported claims

Chunk overlap may preserve boundary context but also duplicates input and does not guarantee complete evidence. See chunking and reranking.

10. Treating prompts or schemas as control boundaries

Gap: “use only the evidence” is treated as guaranteed grounding, or valid JSON as an authorized action.

Repair: make instructions clear and preserve source provenance, then enforce permissions and business rules outside the model. Structured generation can constrain supported output shapes; application validation must still handle refusals, truncation, unsupported schemas and invalid domain values. See structured generation.

Check: a syntactically valid refund for the wrong customer must fail authorization.

Communication pitfalls

11. Monologuing or asking questions without making progress

Give a short roadmap, state bounded assumptions and draw the baseline. Pause at natural decision points so the interviewer can steer. Do not impose an exact interval for checking in or ask for approval after every sentence.

Repair phrase: “The complete request path is here. The main risk is revoked access; I will trace that next unless you want another area first.”

12. Using essay form for parallel requirements

A long paragraph hides missing constraints. Put functional and nonfunctional requirements in separate numbered lists. Use tables for alternatives and failure/repair comparisons, and diagrams for data movement.

Check: can the interviewer point to the requirement that justifies a new queue or model call? See answer frameworks.

13. Explaining one acronym with several more

Start with a standard definition, then a concrete example and the relevant limit. For idempotency, repeated identical requests have the same intended server effect as one request; the responses need not be identical. A refund API can implement this behavior through a scoped operation key and receiver-side deduplication.

A bulkhead reserves separate resources so one workload cannot exhaust all shared capacity. It trades some utilization flexibility for isolation. These definitions are more useful than a sequence of product names.

14. Defending a mistake or agreeing without analysis

If a counterexample exposes an error, state its consequence and repair the relevant decision. If you disagree with a suggestion, explain the requirement or evidence behind the disagreement. Collaboration does not require accepting every proposed component without analysis.

Check: preserve the parts of the design that still satisfy the requirements after a correction.

Interview strategy pitfalls

15. Solving a different problem

A read-only Q&A request does not automatically need autonomous planning or multiple agents. Explain the smallest complete baseline, then add capabilities justified by the scope. Simplicity is useful only while it satisfies the actual requirements.

Check: remove one component and explain which requirement fails. If none does, investigate whether it belongs.

16. Spending the entire session on one deep dive

Use the expected interview length as a planning constraint, while following the interviewer's direction. Reserve room to explain evaluation, recovery, economics and closing decisions. If time becomes short, summarize the remaining risks explicitly rather than claiming they are solved.

Check: can you give a two-minute closing assessment of the design already drawn?

17. Drawing decorative boxes without contracts

Label arrows with the request, event or data they carry. Show state ownership, trust boundaries and the path back to the user. Separate synchronous and asynchronous work. A diagram can be simple visually and still explain sophisticated behavior.

Architecture / visual model
flowchart LR R[Numbered requirement] --> B[Baseline component and contract] B --> F[Specific failure under a changed condition] F --> E[Evidence that locates the failure] E --> C[Repair and added operating cost] C --> V[Verify requirement and remaining limits]
Read diagram source
flowchart LR
    R[Numbered requirement] --> B[Baseline component and contract]
    B --> F[Specific failure under a changed condition]
    F --> E[Evidence that locates the failure]
    E --> C[Repair and added operating cost]
    C --> V[Verify requirement and remaining limits]

Check: trace a timeout and a duplicate request using only the diagram and its contracts.

AI-specific pitfalls

18. Treating inference as a uniform black box

Mistake Concrete failure Mechanism and repair
Equating prefill and decode Long prompts delay short requests Measure prompt processing, decoding and queueing separately; evaluate scheduling changes
Reporting only tokens/second Fast output follows a long initial wait Measure queue time, time to first token, inter-token latency and final-answer time
Sizing from weights alone Model loads but long concurrent requests exhaust memory Add KV cache, activations, runtime workspaces and headroom
Assuming batching lowers every latency Static batch waits violate an interactive deadline Compare throughput and tail latency under the real length distribution
Breaking a cacheable prefix A leading request nonce eliminates reuse Keep stable content first when semantics permit; verify cache rules and isolation
Assuming prompt portability A model swap changes tool choice or format Evaluate the complete model/prompt/tool/retrieval release combination
Trusting successful JSON parsing A missing or renamed field corrupts a business operation Validate versioned schemas, domain rules and caller authorization

At a sustained 40 output tokens/second after the first token, the remaining 199 tokens of a 200-token answer take about 4.975 seconds, in addition to the first-token delay. This is a simplified constant-rate example; real output intervals vary. Define your rate boundary before doing the arithmetic.

See inference fundamentals, KV cache and batching.

19. Treating fluency, citations or confidence as truth

A fluent answer can be wrong. A citation can point to an irrelevant, stale or contradictory passage. A model's stated confidence is not automatically a calibrated probability of correctness.

Use task-specific evidence checks, suitable human review and abstention/escalation behavior. Evaluate these mechanisms rather than promising zero hallucinations. A faithfulness score can pass an answer grounded in an incorrect source. See capability assessment.

20. Leaving security to a final content filter

Authorization, data handling and action permissions belong throughout the architecture. Delimiters and injection detectors can help, but they are not sufficient isolation mechanisms. Tool access, network access and secret exposure determine the consequence of a failure. OWASP prevention guidance.

Check: retrieved content that says “send this file outside the company” must not gain authority from being included in the prompt. See LLM security.

Recovery distinctions worth memorizing

Concept What it establishes What it does not establish
Valid schema Output matches the checked shape The action is permitted or the facts are correct
HTTP success The server reports protocol/application success at that boundary The complete user task is correct
Durable checkpoint Recorded workflow state can be recovered External side effects have been undone
Idempotency key A receiver can identify repeated intended operations under its contract Every service deduplicates forever or an unknown outcome is a failure
Cache hit Stored work was reused Current permission, freshness or correctness
Human approval A person approved the presented action They noticed every error or approved later changes

A worker can crash after a payment commits but before its receipt is saved. Durable execution alone cannot resolve that ambiguity. Reuse the logical operation identity and reconcile with the receiver; distinguish confirmed success, confirmed failure and unknown outcome. See durable execution.

Manager and technical-lead pitfalls

Keep technical depth while explaining delivery:

  1. Name source, application, evaluation, security and incident owners.
  2. Define contracts and service expectations for dependencies owned by other teams.
  3. Stage work around uncertainty: baseline, critical feasibility test, limited rollout, expansion.
  4. State what will be narrowed, delayed or stopped if quality or economics fail.
  5. Describe delegation, coaching and decisions accurately; writing the hardest code is not the only evidence of leadership.

A source owner validates policy content; engineering owns its propagation and enforcement. A release owner makes the exposure decision with evidence. None of these labels excuses an undefined handoff.

Five-minute post-mock review

  1. Write the user outcome and three deciding requirements.
  2. Trace one complete request and one data update.
  3. Identify one severe failure and its detection/recovery path.
  4. Compare your main choice with a feasible alternative.
  5. Check one calculation, including units and missing costs.
  6. State the largest unresolved assumption and who would validate it.

Score each as missing, mentioned or defended with evidence. This is a local practice aid. Then double the load or change a permission and retry the affected explanation.

Interview questions with developed answers

1. Is using a single model automatically a poor design?

No. It may satisfy the requirements with less operating complexity. Add routing only when task-specific evidence and complete economics justify it, including router errors and fallback work.

2. Is filtering after the search engine's candidate stage always a security breach?

No; it depends on the trusted boundaries and the engine's guarantees. The required property is that unauthorized data is not exposed to unauthorized consumers. Explain enforcement inside the trusted service and before text reaches models, users, logs or other unapproved systems.

3. Can a tenant-scoped cache still leak information?

Yes. Users within a tenant may have different access, and permissions may change. Include the relevant security scope and dependencies, and establish current authorization before disclosure.

4. Should missing evidence trigger an unconstrained model answer?

Not when the product promises evidence-grounded answers. Return an explicit limitation, perform permitted bounded retrieval or escalate. A disclaimer does not restore missing evidence.

5. What is wrong with reporting accuracy only on completed calls?

It can hide timeouts and unresolved results. Define the scheduled or eligible population and report completed, failed and unknown outcomes. A conditional score may be useful if clearly labeled alongside coverage.

6. Does a larger context window eliminate retrieval?

No. It changes the feasible input size. Retrieval can still select relevant, current and authorized evidence; direct context can be appropriate for a small packet. Compare both on the task.

7. Does a smaller prompt guarantee that evidence is used correctly?

No. Selection, position, conflicting information and model behavior still matter. Test the particular task instead of treating a token count as a correctness guarantee.

8. Why is 30% lower model-call cost not necessarily 30% lower operating cost?

Other costs may remain unchanged or increase. Include cache writes/misses, tools, review, infrastructure, implementation and operations. Compare equal accepted outcomes and preserve quality requirements.

9. Does idempotency require identical responses?

No. It concerns the intended effect of repeated identical requests. A repeated deletion can return a different response after the resource is already absent while preserving the same intended state.

10. Can rollback to a workflow checkpoint reverse a payment?

No. The external payment has its own state. Reconcile it; any compensation is a separate authorized business operation with its own possible failures.

11. What is missing from a tokens-per-second benchmark?

At least the measurement boundary, prompt/output lengths, concurrency, hardware/runtime, queue delay and quality settings. Interactive experience also depends on first-token and complete-answer latency.

12. Can you fix a retrieval failure by adding a reranker?

Only some failures. A reranker can reorder supplied candidates, but cannot recover a source that never entered that set. First locate the missing evidence in the pipeline.

13. How should you react to a challenge that seems incorrect?

Explain the relevant requirement and evidence respectfully, then test the counterexample. Revise when warranted; do not equate cooperation with uncritical agreement.

14. What should you do when you do not know an SDK guarantee?

State the required behavior and what must be verified, such as deduplication retention and concurrent request semantics. Do not invent a capability from a method name or prior version.

15. What makes a post-mock review useful?

Identify a specific missing causal link and repair it with a definition, calculation, contract or test. Then change a constraint to check whether the explanation still holds.

Final summary and notes

  • Define the task before selecting products.
  • Preserve data meaning, versions and permissions throughout the pipeline.
  • Measure the complete outcome, including failures and uncertainty.
  • Separate correctness, authorization, durability and protocol success.
  • Count human and operating work when comparing cost.
  • Use numbered lists for requirements, tables for choices and labeled diagrams for flows.
  • Finish with the main compromise, owner and next validation.

Continue with answer frameworks and whiteboard exercises.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Answer Frameworks for AI System Design Interviews
NEXT LESSONWhiteboard Exercises for AI System Design →

Explore the diagram