A design pattern is a reusable approach to a recurring problem under stated conditions. It gives you a starting structure and known tradeoffs. It does not establish that your application needs the pattern or that a particular implementation will work.
This Learnastra reference connects requirements to choices you can defend in an interview. Follow a link for the full mechanism, implementation considerations and practice. Reviewed September 24, 2026.
Pattern selection guide
- Define the result: what must the user be able to do, and what would count as success?
- Set constraints: latency, quality, availability, privacy, cost and supported load.
- Draw a baseline: start with the smallest design that can meet the requirements, including a non-AI approach when appropriate.
- Identify a measured or reasoned failure: missing evidence, stale data, queue growth, unsafe action or excessive cost.
- Add one justified mechanism: explain what it fixes, what it costs and what new failure it introduces.
- Close with verification and operations: test the claim, define rollback and name who handles unresolved outcomes.
Read diagram source
flowchart TD
R[Requirements and constraints] --> B[Simplest viable baseline]
B --> E[Measure representative tasks and load]
E --> F{Which constraint fails?}
F -->|Missing evidence| D[Retrieval and data quality]
F -->|Invalid actions| A[Authorization and workflow controls]
F -->|Slow or expensive| C[Capacity, caching or routing]
F -->|No constraint fails| K[Keep baseline and monitor]
D --> V[Evaluate benefit and new risks]
A --> V
C --> V
V --> E
Recall: requirement → baseline → failure → change → evidence. This is a study aid, not a named industry standard.
Retrieval patterns
| Pattern and mechanism | When it helps | Cost or failure to explain |
|---|---|---|
| Basic RAG: retrieve authorized evidence, then generate with it. | Answers need current or private information outside the model. | Bad retrieval, stale evidence or unsupported synthesis can each cause a wrong answer. |
| Hybrid search: combine lexical and vector candidates, then fuse rankings. | Queries mix exact identifiers with paraphrases. | Two indexes, score fusion and update consistency increase operational work. |
| Reranking: rescore a bounded candidate set. | Relevant passages are retrieved but poorly ordered. | Additional inference adds latency; missing candidates remain missing. |
| Query expansion: add alternative expressions or subqueries. | Terminology mismatch limits recall. | Expansion can drift from intent and multiply retrieval calls. |
| HyDE: generate a hypothetical document, then use its representation for retrieval. | A query's form differs substantially from the desired passage. | Invented details can bias retrieval. The hypothetical document is not evidence. |
| Parent–child chunking: retrieve smaller units and supply a related larger passage. | Precise matching needs surrounding context to interpret it. | Larger context can add distraction and cost; parent and child permissions must agree. |
| Graph retrieval: traverse explicit entities and relationships. | Questions require relationship paths or structured joins. | Entity resolution, edge provenance and incremental maintenance can dominate effort. |
| Agentic retrieval: choose subsequent searches based on prior results. | A question needs iterative research or several data sources. | Bound depth, calls and elapsed time; extra searches do not guarantee a complete answer. |
| Abstention and evidence checks: withhold unsupported claims or request clarification. | Missing evidence would make a confident answer harmful. | Excessive abstention reduces usefulness; evaluate answer coverage and correctness together. |
Read diagram source
flowchart LR
Q[Question and authenticated identity] --> P[Resolve scope and policy]
P --> L[Lexical candidates with ACLs]
P --> V[Vector candidates with ACLs]
L --> F[Fuse and deduplicate]
V --> F
F --> R[Optional bounded reranker]
R --> X[Fetch current authorized passages]
X --> G[Generate with provenance]
G --> C[Check citations and claim support]
C --> O[Answer, clarify or abstain]
Interview tip: ask where the failure occurs. Improving generation cannot recover a document that was never indexed. Adding a reranker cannot fix an unauthorized cache hit. Trace ingestion, retrieval and synthesis separately.
Generation patterns
| Pattern and mechanism | When it helps | Cost or failure to explain |
|---|---|---|
| Zero-shot instructions: specify the task without demonstrations. | A clear instruction already meets the quality target. | Measure performance; “simple task” does not prove reliability. |
| Few-shot prompting: include representative input/output demonstrations. | The task needs examples of labels, style or difficult boundaries. | Examples consume context and can overfit a narrow distribution. |
| Reasoning and decomposition: allocate intermediate steps to a difficult problem. | Multi-step problems benefit in task-specific evaluation. | More reasoning can add latency or amplify a wrong premise; APIs differ in supported controls. |
| Self-consistency: sample and aggregate answers. | A checkable task benefits from several candidates. | Correlated errors survive voting. Count all samples, aggregation and verification costs. |
| Structured output: constrain a supported output schema. | Software consumes typed fields rather than free text. | Structure does not prove truth, authorization or business validity; handle refusals and incomplete outputs. |
| Generate then verify: apply a separate check to a proposed result. | A calculator, parser, test suite or evidence checker can test the relevant property. | A second model is not an independent oracle; a weak verifier can accept convincing errors. |
For a current model, confirm supported combinations in the model-selection guide. Do not assume every API accepts temperature, manual reasoning budgets or forced tool selection in every mode.
Agent patterns
| Pattern and mechanism | When it helps | Cost or failure to explain |
|---|---|---|
| ReAct: interleave model decisions, tool execution and observations. | The next step depends on external results. | Runtime budgets, permissions and stop conditions must remain outside the model. |
| Plan and execute: create a plan, run steps and revise when evidence changes. | Dependencies or approval points benefit from an explicit plan. | Plans become stale; a stored plan is not proof that steps completed. |
| Fixed workflow with model steps: code controls transitions around bounded model tasks. | The process has known states and strict business rules. | Less flexibility; exceptional cases need explicit escalation. |
| Multi-agent debate: compare or challenge candidate reasoning. | Diverse candidates improve a measured, verifiable task. | Shared model errors and persuasive but incorrect arguments can dominate. |
| Human approval: pause before a specified consequential action. | A person must accept the exact action or resolve uncertainty. | Queue capacity, stale approvals and reviewer overload can block progress. |
| Handoff: transfer task responsibility to a specialist. | Different tools or expertise justify separate handling. | Define ownership, accepted context, outcome reporting and loop prevention. |
| Advisor / executor: an executor selectively consults another model. | Consultation improves difficult steps without paying that cost on every step. | Frequent consultation can exceed the price and latency of using the stronger model directly. |
| Orchestrator and subagents: delegate bounded work and integrate results. | Work can proceed independently and has a clear merge contract. | Context separation is not credential, filesystem or tenant isolation. Coordination can erase parallel speedups. |
| Durable workflow: persist state and recover across failures and long waits. | Approvals, jobs or tool operations outlive one process. | Runtime recovery does not guarantee exactly-once effects in external services. |
Read diagram source
stateDiagram-v2
[*] --> Ready
Ready --> Decide: budget and deadline available
Decide --> Validate: proposed action
Validate --> AwaitApproval: approval required
Validate --> Escalate: action rejected
AwaitApproval --> Validate: exact action approved
AwaitApproval --> Escalate: denied or expired
Validate --> Execute: authorized, valid and required approval current
Execute --> Record: confirmed outcome
Execute --> Reconcile: timeout or lost response
Reconcile --> Record: outcome resolved
Reconcile --> Escalate: cannot resolve safely
Record --> Ready: more work required
Record --> Complete: completion checks pass
Decide --> Escalate: no safe next step
Ready --> Escalate: budget or deadline exhausted
Complete --> [*]
Escalate --> [*]
For a refund, bind approval to customer, order, amount, currency and operation version. Recheck authorization before execution. An expired approval or changed amount returns for review. An unknown payment result enters reconciliation; it is not silently converted into a second refund request.
Agentic coding patterns
| Pattern | Useful application | Evidence needed before accepting the change |
|---|---|---|
| Read, plan, edit, verify | Refactoring or changing an existing codebase. | Relevant code and repository instructions were inspected; the diff satisfies the task. |
| Scaffold, implement, verify | Building a new feature with known integration points. | The scaffolding works with real dependencies; generated placeholders are resolved. |
| Test-driven change | A behavior can be expressed as a meaningful failing test. | Tests check the required behavior and edge cases, rather than repeating the implementation. |
| Independent diff review | Checking correctness, security and compatibility before merge. | Review findings are validated; neither a second agent nor a clean test run is sufficient alone. |
| Repository instruction files | Supplying architecture context and valid build/check commands. | Use the host's supported file format. Instructions are contextual guidance, not a security sandbox. |
| Bounded parallel work | Independent modules or separate investigation tasks. | Ownership, shared resources, merge order and integration checks are explicit. |
Choose editor, CLI, hosted worker or SDK based on the workflow. The tool comparison explains the boundaries. No single product name establishes autonomy, repeatability or safe execution.
Reliability patterns
| Pattern | Mechanism | Important boundary |
|---|---|---|
| Deadline | Bound the total time allowed across queueing and dependent calls. | Cancelling a client request may not cancel remote computation or a side effect. |
| Retry with backoff and jitter | Retry eligible transient failures with increasing, randomized delays and a maximum attempt count. | Honor provider retry guidance and the remaining deadline. Avoid duplicate effects and retry storms. |
| Circuit breaker | Stop repeatedly calling a dependency during a failure interval; probe recovery. | Separate failure domains and choose meaningful error criteria. |
| Fallback | Switch to another compatible service or a reduced-function response. | Preserve feature, privacy and quality requirements. A second provider may share a regional dependency. |
| Bulkhead and admission control | Bound active work and isolate resource pools. | Queueing without a limit moves overload into memory and latency. |
| Idempotency and reconciliation | Deduplicate operations and resolve ambiguous outcomes using an authoritative record. | Deduplication scope and retention must cover the retry window. |
| Checkpoint and resume | Save sufficient state to continue without losing confirmed progress. | Checkpoint compatibility, nondeterministic steps and external outcomes need explicit handling. |
A deadline example: an interview requires a complete response within 5 seconds. Reserve 0.2 seconds for admission/routing, at most 2.5 seconds for a primary attempt, up to 0.2 seconds for backoff, at most 1.5 seconds for a compatible fallback and 0.3 seconds for final validation/return. That is 4.7 seconds, leaving 0.3 seconds of headroom. These are allocated limits, not a prediction of p95 latency. Shorten or skip a step if the actual remaining deadline is insufficient. Do not give each nested retry a fresh 5-second budget.
Caching patterns
| What is cached | Match or reuse condition | Primary failure to prevent |
|---|---|---|
| Within-request KV state | Compatible prior attention state for the same sequence. | Incorrect positions, model state or memory accounting. |
| Cross-request prefix state | Compatible identical leading tokens, configuration and permitted scope. | Sharing incompatible or private state across isolation boundaries. |
| Exact response | The full answer-affecting key matches and the record is still valid. | Omitting identity, policy, data version, locale or other relevant inputs. |
| Semantic response | A sufficiently equivalent request plus all required scope and validity checks. | Similar questions with different entities, dates, amounts or permissions. |
| Retrieval or embedding result | Matching content/query, model/index version and permissions. | Returning deleted or newly unauthorized content through an old cache. |
Hit rate comes from the workload and key design; cache categories do not have inherent “low,” “medium” or “high” hit rates. A 50% hit rate can be harmful if a small fraction of hits are wrong. Measure valid-hit rate, invalidation lag and the consequence of stale answers.
Tip: start with safely reusable deterministic work when evidence supports it. Semantic response caching is often inappropriate for personal balances, mutable entitlements and irreversible decisions without additional validation.
Security patterns
| Control | What it enforces | What it does not establish |
|---|---|---|
| Authentication and authorization | Identity and permitted operations on resources. | A correct user intent inferred from arbitrary text. |
| Least-privilege tool credentials | Limits on operations and resources a compromised agent can access. | Correct decisions within the permitted scope. |
| Tenant-scoped data paths | Consistent scope through ingestion, retrieval, cache, memory, logs and exports. | Isolation merely because the final search query includes a tenant field. |
| Input and schema validation | Allowed types, ranges and syntactic structure. | Complete prevention of prompt injection or factual errors. |
| Untrusted-content separation | Clear provenance and reduced opportunity for external text to be treated as application instructions. | A guarantee that the model will never follow malicious content. |
| Sandbox and egress policy | Runtime restrictions on files, processes, network destinations and resources. | A claim that every container configuration is a strong hostile-tenant boundary. |
| Output and disclosure controls | Checks for prohibited disclosures or unsafe output handling. | Recovery of data already sent to an unauthorized provider or tool. |
| Quotas and budget reservations | Limits on concurrent and cumulative resource use. | A strict cap if concurrent calls only read the budget and reserve nothing atomically. |
Evaluation patterns
| Pattern | What it answers | Required qualification |
|---|---|---|
| Versioned regression set | Does the candidate retain required behavior on known cases? | Maintain held-out cases; a development set repeatedly optimized against is not an independent test. |
| Deterministic verifier | Does a checkable contract hold, such as schema or expected state? | Incomplete tests can miss important errors. |
| Calibrated model judge | How does an output meet a defined semantic rubric? | Compare with expert labels; track judge failures, bias and unknown outcomes. |
| Expert review | How do qualified reviewers assess difficult cases? | Reviewers can disagree; document the rubric, adjudication and uncertainty. |
| Paired offline comparison | How do two candidates perform on the same cases? | Preserve pairing and report uncertainty, important slices and regression counts. |
| Online experiment | Does a change improve a production outcome under random assignment? | Define the randomization unit, interference risks, stopping rule and guardrail metrics. |
| Load and failure testing | What happens under concurrency, dependency failure and recovery? | Token and request quotas, queues and human review capacity all matter. |
Cost optimization patterns
| Candidate change | Source of possible benefit | Additional cost or risk |
|---|---|---|
| Model routing | Use less costly candidates where quality remains adequate. | Router calls, extra attempts, feature mismatches and human corrections. |
| Caching | Reuse computation or valid outputs. | Storage, writes, invalidation, isolation and stale-answer handling. |
| Context reduction | Send fewer unnecessary tokens. | Removing a crucial exception or citation can increase errors and rework. |
| Offline batch execution | Use asynchronous capacity and applicable provider discounts. | Completion windows, lifecycle support, retries and result retention. |
| Distillation or fine-tuning | Make a narrower model serve a stable, high-volume task. | Data rights, training, evaluation, serving, maintenance and quality loss. |
Worked decision: cheaper model routing
Functional requirements
- Classify and answer 100,000 support requests per month.
- Escalate cases outside the model's approved scope to a person.
- Preserve the conversation and resolution outcome for review.
Non-functional requirements
- Meet the same task-quality and data-handling criteria as the baseline.
- Keep the agreed interactive latency target on each important traffic slice.
- Reduce total monthly cost without exceeding available reviewer capacity.
Baseline: use one qualified model for every request, then review 2% of requests. Candidate: a router sends 60% to a smaller model and 40% to the baseline model. The following are hypothetical loaded costs, not provider prices or measured outcomes.
| Monthly cost | Baseline | Routing candidate |
|---|---|---|
| Generation | 100,000 × $0.08 = $8,000 | 60,000 × $0.02 + 40,000 × $0.08 = $4,400 |
| Router | $0 | 100,000 × $0.002 = $200 |
| Human review at 4 minutes and $45/hour | 2,000 × $3 = $6,000 | 3,000 × $3 = $9,000 |
| Operations allocation | $1,200 | $1,800 |
| Change implementation amortization | $0 | $600 |
| Common infrastructure/support | $2,000 | $2,000 |
| Total | $17,200 | $18,000 |
The flaw: model and router spend falls by $3,400, but additional review and operating costs make the candidate $800 more expensive. Review workload rises from about 133.3 to 200 hours per month. The team must account for that additional capacity even if headcount cannot change immediately.
Repair: restrict small-model routing to validated categories. Re-evaluate quality, traffic shares, review rate and latency; do not assume the original 60/40 traffic split survives. If that split does remain and review returns to 2%, the candidate totals $15,000, saving $2,200 per month. With all other assumptions fixed, break-even review workload is about 2,733 requests, or 2.73%. Any routing-induced harm not captured by review also belongs in the decision.
Closing: proceed only if a held-out evaluation and limited rollout support the revised assumptions. Retain the baseline route for rollback, log the policy version and compare cost per correctly resolved request. A lower token bill is evidence about one line item, not the final business case.
Anti-patterns to avoid
| Tempting shortcut | Why it fails | Better interview answer |
|---|---|---|
| Add RAG to every AI product. | Some tasks need classification, calculation or a structured database query. | Select evidence access from the task's information requirement. |
| Put all available material in context. | Limits, distraction, permissions and stale information remain. | Select useful authorized evidence and measure lost-answer cases. |
| Retry until it works. | Amplifies load and can duplicate effects. | Bound attempts and deadlines; reconcile ambiguous operations. |
| Add another provider for guaranteed availability. | Dependencies can correlate and contracts can differ. | Evaluate a compatible fallback against a simpler degraded response. |
| Trust output because two models agree. | Their errors or sources can be shared. | Verify the required property using suitable evidence. |
| Trace everything forever. | Sensitive data, retention and storage costs accumulate. | Record the minimum useful telemetry under explicit access and retention rules. |
| Add a critic to stop infinite loops. | A critic can also fail to terminate or recognize progress. | Enforce runtime budgets and stop conditions independently. |
| Accept a screenshot as proof a payment completed. | The UI can be stale or the response lost. | Confirm through authoritative state; reconcile uncertainty. |
| Treat prompt instructions as security policy. | The model may ignore them or read hostile instructions. | Enforce credentials, authorization and execution restrictions. |
| Turn reasoning off for every easy task. | Some models do not expose that mode; classification itself can fail. | Choose supported controls and validate the complete routing policy. |
| Start with semantic caching because it sounds cheap. | Similarity can match the wrong account, date or intention. | First establish safe reuse conditions and measure valid hits. |
Interview practice
- Reranker or hybrid retrieval? A relevant part number never appears in the vector candidate set. Start by improving candidate recall, for example lexical retrieval and fusion; a reranker cannot score a missing candidate.
- RAG or fine-tuning? A policy changes daily. Retrieve a versioned authoritative policy; fine-tuning is not a dependable daily fact store.
- Agent or workflow? Every invoice follows four known approval states. A workflow with bounded model extraction is a strong baseline; justify any model-selected transitions.
- Retry or reconcile? A refund request times out after transmission. Resolve its status using the same operation identity before initiating another effect.
- One model or two? The current model already meets quality, latency and availability targets. A second model needs evidence of incremental benefit exceeding cost and complexity.
- Cache or recompute? Two users ask the same question but have different permissions. Similar text is insufficient; use permission-scoped validity checks or recompute.
- Parallel agents or sequential work? Two changes modify the same schema and depend on its final form. Resolve that dependency and ownership first; parallel work is not automatically independent.
- Human approval or automation? An action is approved, then its amount changes. Invalidate approval and seek authorization for the actual action.
- More samples or a better verifier? Five answers repeat the same unsupported claim. Inspect shared evidence and the acceptance test before buying more votes.
- Higher throughput or lower latency? A larger batch improves GPU utilization but grows the queue. Measure the end-to-end latency distribution under the arrival pattern, not only kernel throughput.
- Cheaper inference or cheaper service? Inference savings cause more correction work. Include review, rework, incidents and operations in the comparison.
- Better benchmark or better product? A public score rises while domain performance falls. Prefer the application's representative held-out tasks and required slices for that deployment decision.
- More logging or better debugging? Full prompts expose customer secrets. Preserve scoped identifiers, versions, timings and safely retained samples needed for diagnosis.
- Longer context or better retrieval? Relevant evidence is buried among distractors. Compare both on matched tasks and budgets; capacity alone is not evidence-use quality.
- Timeout or cancellation? The caller returns an error while a worker continues. Persist job state, stop further scheduling where possible and reconcile any external action already issued.
Final summary and notes
| If the problem is… | First investigate… | Then consider… |
|---|---|---|
| Wrong evidence | Ingestion, authorization, labels and candidate recall | Hybrid retrieval, chunk changes or reranking |
| Wrong reasoning | Task specification, evidence and a valid outcome check | Different model, decomposition or verified sampling |
| Unsafe action | Identity, permissions and operation state | Approval, scoped tools and sandbox controls |
| Slow service | Queueing and the measured critical path | Admission limits, capacity, batching or reduced work |
| High cost | Full cost per correct outcome | Routing, caching, context reduction or adaptation |
| Unreliable recovery | Recorded state and unknown external outcomes | Idempotency, reconciliation and durable execution |
End an interview by naming the chosen design, the assumption most likely to fail, the metric that reveals it and the safe fallback. Continue with the complete pattern lesson, anti-pattern lesson or whiteboard interviews.