Learnastra AI SYSTEM DESIGNAnup Rai

Interview toolkit

AI Architecture Pattern Reference

By Anup Rai15 min readReviewed September 2026

A design pattern is a reusable approach to a recurring problem under stated conditions. It gives you a starting structure and known tradeoffs. It does not establish that your application needs the pattern or that a particular implementation will work.

This Learnastra reference connects requirements to choices you can defend in an interview. Follow a link for the full mechanism, implementation considerations and practice. Reviewed September 24, 2026.

Pattern selection guide

  1. Define the result: what must the user be able to do, and what would count as success?
  2. Set constraints: latency, quality, availability, privacy, cost and supported load.
  3. Draw a baseline: start with the smallest design that can meet the requirements, including a non-AI approach when appropriate.
  4. Identify a measured or reasoned failure: missing evidence, stale data, queue growth, unsafe action or excessive cost.
  5. Add one justified mechanism: explain what it fixes, what it costs and what new failure it introduces.
  6. Close with verification and operations: test the claim, define rollback and name who handles unresolved outcomes.
Architecture / visual model
flowchart TD R[Requirements and constraints] --> B[Simplest viable baseline] B --> E[Measure representative tasks and load] E --> F{Which constraint fails?} F -->|Missing evidence| D[Retrieval and data quality] F -->|Invalid actions| A[Authorization and workflow controls] F -->|Slow or expensive| C[Capacity, caching or routing] F -->|No constraint fails| K[Keep baseline and monitor] D --> V[Evaluate benefit and new risks] A --> V C --> V V --> E
Read diagram source
flowchart TD
  R[Requirements and constraints] --> B[Simplest viable baseline]
  B --> E[Measure representative tasks and load]
  E --> F{Which constraint fails?}
  F -->|Missing evidence| D[Retrieval and data quality]
  F -->|Invalid actions| A[Authorization and workflow controls]
  F -->|Slow or expensive| C[Capacity, caching or routing]
  F -->|No constraint fails| K[Keep baseline and monitor]
  D --> V[Evaluate benefit and new risks]
  A --> V
  C --> V
  V --> E

Recall: requirement → baseline → failure → change → evidence. This is a study aid, not a named industry standard.

Retrieval patterns

Pattern and mechanism When it helps Cost or failure to explain
Basic RAG: retrieve authorized evidence, then generate with it. Answers need current or private information outside the model. Bad retrieval, stale evidence or unsupported synthesis can each cause a wrong answer.
Hybrid search: combine lexical and vector candidates, then fuse rankings. Queries mix exact identifiers with paraphrases. Two indexes, score fusion and update consistency increase operational work.
Reranking: rescore a bounded candidate set. Relevant passages are retrieved but poorly ordered. Additional inference adds latency; missing candidates remain missing.
Query expansion: add alternative expressions or subqueries. Terminology mismatch limits recall. Expansion can drift from intent and multiply retrieval calls.
HyDE: generate a hypothetical document, then use its representation for retrieval. A query's form differs substantially from the desired passage. Invented details can bias retrieval. The hypothetical document is not evidence.
Parent–child chunking: retrieve smaller units and supply a related larger passage. Precise matching needs surrounding context to interpret it. Larger context can add distraction and cost; parent and child permissions must agree.
Graph retrieval: traverse explicit entities and relationships. Questions require relationship paths or structured joins. Entity resolution, edge provenance and incremental maintenance can dominate effort.
Agentic retrieval: choose subsequent searches based on prior results. A question needs iterative research or several data sources. Bound depth, calls and elapsed time; extra searches do not guarantee a complete answer.
Abstention and evidence checks: withhold unsupported claims or request clarification. Missing evidence would make a confident answer harmful. Excessive abstention reduces usefulness; evaluate answer coverage and correctness together.
Architecture / visual model
flowchart LR Q[Question and authenticated identity] --> P[Resolve scope and policy] P --> L[Lexical candidates with ACLs] P --> V[Vector candidates with ACLs] L --> F[Fuse and deduplicate] V --> F F --> R[Optional bounded reranker] R --> X[Fetch current authorized passages] X --> G[Generate with provenance] G --> C[Check citations and claim support] C --> O[Answer, clarify or abstain]
Read diagram source
flowchart LR
  Q[Question and authenticated identity] --> P[Resolve scope and policy]
  P --> L[Lexical candidates with ACLs]
  P --> V[Vector candidates with ACLs]
  L --> F[Fuse and deduplicate]
  V --> F
  F --> R[Optional bounded reranker]
  R --> X[Fetch current authorized passages]
  X --> G[Generate with provenance]
  G --> C[Check citations and claim support]
  C --> O[Answer, clarify or abstain]

Interview tip: ask where the failure occurs. Improving generation cannot recover a document that was never indexed. Adding a reranker cannot fix an unauthorized cache hit. Trace ingestion, retrieval and synthesis separately.

Generation patterns

Pattern and mechanism When it helps Cost or failure to explain
Zero-shot instructions: specify the task without demonstrations. A clear instruction already meets the quality target. Measure performance; “simple task” does not prove reliability.
Few-shot prompting: include representative input/output demonstrations. The task needs examples of labels, style or difficult boundaries. Examples consume context and can overfit a narrow distribution.
Reasoning and decomposition: allocate intermediate steps to a difficult problem. Multi-step problems benefit in task-specific evaluation. More reasoning can add latency or amplify a wrong premise; APIs differ in supported controls.
Self-consistency: sample and aggregate answers. A checkable task benefits from several candidates. Correlated errors survive voting. Count all samples, aggregation and verification costs.
Structured output: constrain a supported output schema. Software consumes typed fields rather than free text. Structure does not prove truth, authorization or business validity; handle refusals and incomplete outputs.
Generate then verify: apply a separate check to a proposed result. A calculator, parser, test suite or evidence checker can test the relevant property. A second model is not an independent oracle; a weak verifier can accept convincing errors.

For a current model, confirm supported combinations in the model-selection guide. Do not assume every API accepts temperature, manual reasoning budgets or forced tool selection in every mode.

Agent patterns

Pattern and mechanism When it helps Cost or failure to explain
ReAct: interleave model decisions, tool execution and observations. The next step depends on external results. Runtime budgets, permissions and stop conditions must remain outside the model.
Plan and execute: create a plan, run steps and revise when evidence changes. Dependencies or approval points benefit from an explicit plan. Plans become stale; a stored plan is not proof that steps completed.
Fixed workflow with model steps: code controls transitions around bounded model tasks. The process has known states and strict business rules. Less flexibility; exceptional cases need explicit escalation.
Multi-agent debate: compare or challenge candidate reasoning. Diverse candidates improve a measured, verifiable task. Shared model errors and persuasive but incorrect arguments can dominate.
Human approval: pause before a specified consequential action. A person must accept the exact action or resolve uncertainty. Queue capacity, stale approvals and reviewer overload can block progress.
Handoff: transfer task responsibility to a specialist. Different tools or expertise justify separate handling. Define ownership, accepted context, outcome reporting and loop prevention.
Advisor / executor: an executor selectively consults another model. Consultation improves difficult steps without paying that cost on every step. Frequent consultation can exceed the price and latency of using the stronger model directly.
Orchestrator and subagents: delegate bounded work and integrate results. Work can proceed independently and has a clear merge contract. Context separation is not credential, filesystem or tenant isolation. Coordination can erase parallel speedups.
Durable workflow: persist state and recover across failures and long waits. Approvals, jobs or tool operations outlive one process. Runtime recovery does not guarantee exactly-once effects in external services.
Architecture / visual model
stateDiagram-v2 [*] --> Ready Ready --> Decide: budget and deadline available Decide --> Validate: proposed action Validate --> AwaitApproval: approval required Validate --> Escalate: action rejected AwaitApproval --> Validate: exact action approved AwaitApproval --> Escalate: denied or expired Validate --> Execute: authorized, valid and required approval current Execute --> Record: confirmed outcome Execute --> Reconcile: timeout or lost response Reconcile --> Record: outcome resolved Reconcile --> Escalate: cannot resolve safely Record --> Ready: more work required Record --> Complete: completion checks pass Decide --> Escalate: no safe next step Ready --> Escalate: budget or deadline exhausted Complete --> [*] Escalate --> [*]
Read diagram source
stateDiagram-v2
  [*] --> Ready
  Ready --> Decide: budget and deadline available
  Decide --> Validate: proposed action
  Validate --> AwaitApproval: approval required
  Validate --> Escalate: action rejected
  AwaitApproval --> Validate: exact action approved
  AwaitApproval --> Escalate: denied or expired
  Validate --> Execute: authorized, valid and required approval current
  Execute --> Record: confirmed outcome
  Execute --> Reconcile: timeout or lost response
  Reconcile --> Record: outcome resolved
  Reconcile --> Escalate: cannot resolve safely
  Record --> Ready: more work required
  Record --> Complete: completion checks pass
  Decide --> Escalate: no safe next step
  Ready --> Escalate: budget or deadline exhausted
  Complete --> [*]
  Escalate --> [*]

For a refund, bind approval to customer, order, amount, currency and operation version. Recheck authorization before execution. An expired approval or changed amount returns for review. An unknown payment result enters reconciliation; it is not silently converted into a second refund request.

Agentic coding patterns

Pattern Useful application Evidence needed before accepting the change
Read, plan, edit, verify Refactoring or changing an existing codebase. Relevant code and repository instructions were inspected; the diff satisfies the task.
Scaffold, implement, verify Building a new feature with known integration points. The scaffolding works with real dependencies; generated placeholders are resolved.
Test-driven change A behavior can be expressed as a meaningful failing test. Tests check the required behavior and edge cases, rather than repeating the implementation.
Independent diff review Checking correctness, security and compatibility before merge. Review findings are validated; neither a second agent nor a clean test run is sufficient alone.
Repository instruction files Supplying architecture context and valid build/check commands. Use the host's supported file format. Instructions are contextual guidance, not a security sandbox.
Bounded parallel work Independent modules or separate investigation tasks. Ownership, shared resources, merge order and integration checks are explicit.

Choose editor, CLI, hosted worker or SDK based on the workflow. The tool comparison explains the boundaries. No single product name establishes autonomy, repeatability or safe execution.

Reliability patterns

Pattern Mechanism Important boundary
Deadline Bound the total time allowed across queueing and dependent calls. Cancelling a client request may not cancel remote computation or a side effect.
Retry with backoff and jitter Retry eligible transient failures with increasing, randomized delays and a maximum attempt count. Honor provider retry guidance and the remaining deadline. Avoid duplicate effects and retry storms.
Circuit breaker Stop repeatedly calling a dependency during a failure interval; probe recovery. Separate failure domains and choose meaningful error criteria.
Fallback Switch to another compatible service or a reduced-function response. Preserve feature, privacy and quality requirements. A second provider may share a regional dependency.
Bulkhead and admission control Bound active work and isolate resource pools. Queueing without a limit moves overload into memory and latency.
Idempotency and reconciliation Deduplicate operations and resolve ambiguous outcomes using an authoritative record. Deduplication scope and retention must cover the retry window.
Checkpoint and resume Save sufficient state to continue without losing confirmed progress. Checkpoint compatibility, nondeterministic steps and external outcomes need explicit handling.

A deadline example: an interview requires a complete response within 5 seconds. Reserve 0.2 seconds for admission/routing, at most 2.5 seconds for a primary attempt, up to 0.2 seconds for backoff, at most 1.5 seconds for a compatible fallback and 0.3 seconds for final validation/return. That is 4.7 seconds, leaving 0.3 seconds of headroom. These are allocated limits, not a prediction of p95 latency. Shorten or skip a step if the actual remaining deadline is insufficient. Do not give each nested retry a fresh 5-second budget.

Caching patterns

What is cached Match or reuse condition Primary failure to prevent
Within-request KV state Compatible prior attention state for the same sequence. Incorrect positions, model state or memory accounting.
Cross-request prefix state Compatible identical leading tokens, configuration and permitted scope. Sharing incompatible or private state across isolation boundaries.
Exact response The full answer-affecting key matches and the record is still valid. Omitting identity, policy, data version, locale or other relevant inputs.
Semantic response A sufficiently equivalent request plus all required scope and validity checks. Similar questions with different entities, dates, amounts or permissions.
Retrieval or embedding result Matching content/query, model/index version and permissions. Returning deleted or newly unauthorized content through an old cache.

Hit rate comes from the workload and key design; cache categories do not have inherent “low,” “medium” or “high” hit rates. A 50% hit rate can be harmful if a small fraction of hits are wrong. Measure valid-hit rate, invalidation lag and the consequence of stale answers.

Tip: start with safely reusable deterministic work when evidence supports it. Semantic response caching is often inappropriate for personal balances, mutable entitlements and irreversible decisions without additional validation.

Security patterns

Control What it enforces What it does not establish
Authentication and authorization Identity and permitted operations on resources. A correct user intent inferred from arbitrary text.
Least-privilege tool credentials Limits on operations and resources a compromised agent can access. Correct decisions within the permitted scope.
Tenant-scoped data paths Consistent scope through ingestion, retrieval, cache, memory, logs and exports. Isolation merely because the final search query includes a tenant field.
Input and schema validation Allowed types, ranges and syntactic structure. Complete prevention of prompt injection or factual errors.
Untrusted-content separation Clear provenance and reduced opportunity for external text to be treated as application instructions. A guarantee that the model will never follow malicious content.
Sandbox and egress policy Runtime restrictions on files, processes, network destinations and resources. A claim that every container configuration is a strong hostile-tenant boundary.
Output and disclosure controls Checks for prohibited disclosures or unsafe output handling. Recovery of data already sent to an unauthorized provider or tool.
Quotas and budget reservations Limits on concurrent and cumulative resource use. A strict cap if concurrent calls only read the budget and reserve nothing atomically.

Evaluation patterns

Pattern What it answers Required qualification
Versioned regression set Does the candidate retain required behavior on known cases? Maintain held-out cases; a development set repeatedly optimized against is not an independent test.
Deterministic verifier Does a checkable contract hold, such as schema or expected state? Incomplete tests can miss important errors.
Calibrated model judge How does an output meet a defined semantic rubric? Compare with expert labels; track judge failures, bias and unknown outcomes.
Expert review How do qualified reviewers assess difficult cases? Reviewers can disagree; document the rubric, adjudication and uncertainty.
Paired offline comparison How do two candidates perform on the same cases? Preserve pairing and report uncertainty, important slices and regression counts.
Online experiment Does a change improve a production outcome under random assignment? Define the randomization unit, interference risks, stopping rule and guardrail metrics.
Load and failure testing What happens under concurrency, dependency failure and recovery? Token and request quotas, queues and human review capacity all matter.

Cost optimization patterns

Candidate change Source of possible benefit Additional cost or risk
Model routing Use less costly candidates where quality remains adequate. Router calls, extra attempts, feature mismatches and human corrections.
Caching Reuse computation or valid outputs. Storage, writes, invalidation, isolation and stale-answer handling.
Context reduction Send fewer unnecessary tokens. Removing a crucial exception or citation can increase errors and rework.
Offline batch execution Use asynchronous capacity and applicable provider discounts. Completion windows, lifecycle support, retries and result retention.
Distillation or fine-tuning Make a narrower model serve a stable, high-volume task. Data rights, training, evaluation, serving, maintenance and quality loss.

Worked decision: cheaper model routing

Functional requirements

  1. Classify and answer 100,000 support requests per month.
  2. Escalate cases outside the model's approved scope to a person.
  3. Preserve the conversation and resolution outcome for review.

Non-functional requirements

  1. Meet the same task-quality and data-handling criteria as the baseline.
  2. Keep the agreed interactive latency target on each important traffic slice.
  3. Reduce total monthly cost without exceeding available reviewer capacity.

Baseline: use one qualified model for every request, then review 2% of requests. Candidate: a router sends 60% to a smaller model and 40% to the baseline model. The following are hypothetical loaded costs, not provider prices or measured outcomes.

Monthly cost Baseline Routing candidate
Generation 100,000 × $0.08 = $8,000 60,000 × $0.02 + 40,000 × $0.08 = $4,400
Router $0 100,000 × $0.002 = $200
Human review at 4 minutes and $45/hour 2,000 × $3 = $6,000 3,000 × $3 = $9,000
Operations allocation $1,200 $1,800
Change implementation amortization $0 $600
Common infrastructure/support $2,000 $2,000
Total $17,200 $18,000

The flaw: model and router spend falls by $3,400, but additional review and operating costs make the candidate $800 more expensive. Review workload rises from about 133.3 to 200 hours per month. The team must account for that additional capacity even if headcount cannot change immediately.

Repair: restrict small-model routing to validated categories. Re-evaluate quality, traffic shares, review rate and latency; do not assume the original 60/40 traffic split survives. If that split does remain and review returns to 2%, the candidate totals $15,000, saving $2,200 per month. With all other assumptions fixed, break-even review workload is about 2,733 requests, or 2.73%. Any routing-induced harm not captured by review also belongs in the decision.

Closing: proceed only if a held-out evaluation and limited rollout support the revised assumptions. Retain the baseline route for rollback, log the policy version and compare cost per correctly resolved request. A lower token bill is evidence about one line item, not the final business case.

Anti-patterns to avoid

Tempting shortcut Why it fails Better interview answer
Add RAG to every AI product. Some tasks need classification, calculation or a structured database query. Select evidence access from the task's information requirement.
Put all available material in context. Limits, distraction, permissions and stale information remain. Select useful authorized evidence and measure lost-answer cases.
Retry until it works. Amplifies load and can duplicate effects. Bound attempts and deadlines; reconcile ambiguous operations.
Add another provider for guaranteed availability. Dependencies can correlate and contracts can differ. Evaluate a compatible fallback against a simpler degraded response.
Trust output because two models agree. Their errors or sources can be shared. Verify the required property using suitable evidence.
Trace everything forever. Sensitive data, retention and storage costs accumulate. Record the minimum useful telemetry under explicit access and retention rules.
Add a critic to stop infinite loops. A critic can also fail to terminate or recognize progress. Enforce runtime budgets and stop conditions independently.
Accept a screenshot as proof a payment completed. The UI can be stale or the response lost. Confirm through authoritative state; reconcile uncertainty.
Treat prompt instructions as security policy. The model may ignore them or read hostile instructions. Enforce credentials, authorization and execution restrictions.
Turn reasoning off for every easy task. Some models do not expose that mode; classification itself can fail. Choose supported controls and validate the complete routing policy.
Start with semantic caching because it sounds cheap. Similarity can match the wrong account, date or intention. First establish safe reuse conditions and measure valid hits.

Interview practice

  1. Reranker or hybrid retrieval? A relevant part number never appears in the vector candidate set. Start by improving candidate recall, for example lexical retrieval and fusion; a reranker cannot score a missing candidate.
  2. RAG or fine-tuning? A policy changes daily. Retrieve a versioned authoritative policy; fine-tuning is not a dependable daily fact store.
  3. Agent or workflow? Every invoice follows four known approval states. A workflow with bounded model extraction is a strong baseline; justify any model-selected transitions.
  4. Retry or reconcile? A refund request times out after transmission. Resolve its status using the same operation identity before initiating another effect.
  5. One model or two? The current model already meets quality, latency and availability targets. A second model needs evidence of incremental benefit exceeding cost and complexity.
  6. Cache or recompute? Two users ask the same question but have different permissions. Similar text is insufficient; use permission-scoped validity checks or recompute.
  7. Parallel agents or sequential work? Two changes modify the same schema and depend on its final form. Resolve that dependency and ownership first; parallel work is not automatically independent.
  8. Human approval or automation? An action is approved, then its amount changes. Invalidate approval and seek authorization for the actual action.
  9. More samples or a better verifier? Five answers repeat the same unsupported claim. Inspect shared evidence and the acceptance test before buying more votes.
  10. Higher throughput or lower latency? A larger batch improves GPU utilization but grows the queue. Measure the end-to-end latency distribution under the arrival pattern, not only kernel throughput.
  11. Cheaper inference or cheaper service? Inference savings cause more correction work. Include review, rework, incidents and operations in the comparison.
  12. Better benchmark or better product? A public score rises while domain performance falls. Prefer the application's representative held-out tasks and required slices for that deployment decision.
  13. More logging or better debugging? Full prompts expose customer secrets. Preserve scoped identifiers, versions, timings and safely retained samples needed for diagnosis.
  14. Longer context or better retrieval? Relevant evidence is buried among distractors. Compare both on matched tasks and budgets; capacity alone is not evidence-use quality.
  15. Timeout or cancellation? The caller returns an error while a worker continues. Persist job state, stop further scheduling where possible and reconcile any external action already issued.

Final summary and notes

If the problem is… First investigate… Then consider…
Wrong evidence Ingestion, authorization, labels and candidate recall Hybrid retrieval, chunk changes or reranking
Wrong reasoning Task specification, evidence and a valid outcome check Different model, decomposition or verified sampling
Unsafe action Identity, permissions and operation state Approval, scoped tools and sandbox controls
Slow service Queueing and the measured critical path Admission limits, capacity, batching or reduced work
High cost Full cost per correct outcome Routing, caching, context reduction or adaptation
Unreliable recovery Recorded state and unknown external outcomes Idempotency, reconciliation and durable execution

End an interview by naming the chosen design, the assumption most likely to fail, the metric that reveals it and the safe fallback. Continue with the complete pattern lesson, anti-pattern lesson or whiteboard interviews.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Learnastra AI Interview Guide
NEXT LESSONAI Engineering Glossary →

Explore the diagram