Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Context Engineering

By Anup Rai9 min readReviewed September 2026

Numerical examples are illustrative unless explicitly sourced.

Remember: Keep the task, evidence, and state that the next decision needs.

Definition and scope

Context engineering is the design of how information is selected, organized and supplied to a model for each call. That information can include instructions, the current request, conversation history, tool definitions, retrieved evidence and previous tool results. The goal is to provide sufficient relevant information within the task's quality, latency, cost and access constraints.

Term What it holds Example in a refund workflow
Context Inputs actually supplied to one model call Current request and selected order evidence
Memory Stored information available for later retrieval A retained communication preference
Business state Authoritative facts about the domain Payment and refund records
Durable execution state Workflow progress and recovery information An operation ID and whether its outcome is known

A record in storage is not available to the model unless the application loads it. A model-generated summary is not the authority for whether a refund was approved or completed.

Assemble the next call step by step

  1. Load governing instructions and the current user request.
  2. Resolve relevant business and workflow state through the authenticated application.
  3. Select permitted evidence with source, version and retrieval-time references.
  4. Include the tool schemas needed for the next step and the recent interaction needed to interpret it.
  5. Reserve output capacity, serialize the actual request and count its tokens.
  6. Record selected source IDs and context-policy version so failures can be diagnosed.
Architecture / visual model
flowchart LR U[Current task and instructions] --> C[Context assembly] S[Authoritative state] --> C R[Permitted retrieved evidence] --> C H[Selected history and notes] --> C C --> B[Token budget and provenance checks] B --> M[Next model call] M --> V[Validate proposal against current state]
Read diagram source
flowchart LR
    U[Current task and instructions] --> C[Context assembly]
    S[Authoritative state] --> C
    R[Permitted retrieved evidence] --> C
    H[Selected history and notes] --> C
    C --> B[Token budget and provenance checks]
    B --> M[Next model call]
    M --> V[Validate proposal against current state]

Distinguish selection from compaction. Selection omits material that is not needed now but can be fetched later. Compaction replaces a longer history with a shorter representation and can lose or distort information. When compacting, preserve commitments, unresolved questions, completed and uncertain actions, and references to authoritative records. Do not allow a summary to transform “consider this booking” into “booking approved.”

A conceptual state record might be:

Goal: assess a duplicate-charge refund request
Constraints: keep the subscription active
Evidence: order and payment record IDs, with versions
Decisions: refund proposed; no approval recorded
Open questions: confirm whether the second charge settled
Actions: payment lookup complete; refund not attempted

The record is not the model's hidden reasoning. It is application-visible information designed to support the next step. Validate consequential facts against their authoritative sources, especially after long waits or permission changes.

Five practical techniques

Technique Plain meaning Main tradeoff
Selection Keep relevant evidence and discard noise Missing a crucial qualification
Just-in-time loading Carry references, fetch details when needed Extra calls and latency
Compaction Summarize older interaction Omission or distortion
Structured notes Persist goals, decisions, and unresolved work Stale or poisoned memory
Task isolation Give a bounded subtask its own context Handoff loss and coordination cost

These techniques are described in Anthropic's context-engineering guidance. They are options to measure, not instructions to add multiple agents to every system.

What a summary must preserve

For a refund workflow, preserve user intent, exact approved action, evidence references, unresolved questions, operation IDs, and known outcomes. Store authoritative payment and approval records separately. A generated summary cannot turn “refund proposed” into “refund approved.” Recheck current permissions before using retrieved or remembered data.

Keep access to the underlying evidence so a consequential claim can be verified. Version summaries and retain provenance where needed. Treat notes derived from untrusted documents as untrusted evidence; persistence does not promote them to policy.

Event Preserve in context or notes Verify outside the summary
User corrects the amount The correction and which proposal it supersedes The new exact amount before approval
Tool times out after submission Operation ID and unknown outcome Provider status before retrying
Access is revoked The fact that current access needs checking Live authorization on retrieval/execution
Two records conflict Both source IDs and the unresolved conflict Authoritative record and reconciliation result
External text claims approval Its source and untrusted status Actual approval record bound to the action

Long context versus retrieval

Document count is not a capacity measure: ten giant PDFs can exceed a budget while thousands of tiny records may not. Count with the actual tokenizer and include tool schemas, modalities, history, and output allocation.

Direct context is attractive when the supplied corpus is bounded, relevant, and permitted, especially for cross-document work. Retrieval helps select from larger or changing corpora. A hybrid can load a small stable core and retrieve the rest. Compare answer quality, useful evidence coverage, freshness, permission handling, cost, and complete-request latency.

A document being present in the prompt does not guarantee the model will use it correctly. Position, distractors, contradictions, and task complexity matter. Lost in the Middle demonstrates position sensitivity in the studied models; do not assume every newer model has the same curve or that the issue has disappeared.

Budget and reasoning controls

Reserve enough room for the requested response and model-specific reasoning accounting. Exact context/output limits and supported effort settings depend on the model/API. More reasoning is not a guaranteed monotonic quality improvement, and some models do not allow reasoning to be disabled.

Measure configuration changes on your workload. Avoid fixed claims such as “high effort costs 20 times more” or “a confidence score below 0.5 means no reasoning is needed.” The router and its mistakes must be evaluated too. Current model-specific constraints belong in the model selection guide.

Caching does not make context free

Provider prompt caching can reuse eligible prefix processing. Rates, write fees, TTL, minimum sizes, eviction, and supported prefixes vary. Even a cache hit may have token charges and decode attention costs. It does not guarantee the latency of a tiny prompt or a 100% hit rate.

Compare total workload cost including cache creation/storage and invalidation with retrieval cost. “Two reuses beat RAG” is not a universal break-even rule. Answer caching is different: it reuses the response itself and requires correctness, permission, and freshness checks.

Diagnose long-session failure

Compare short and long task traces. Look for lost constraints, repeated irrelevant tool output, stale summaries, conflicting instructions, and missing evidence. Change one context policy and measure completion, severe errors, token use, tool calls, and latency. Do not attribute every failure to quadratic attention; model behavior and system design have multiple causes.

Manager follow-ups

“What is the first launch version?” Explicit context assembly, size limits, useful evidence, and observable failures before complex memory automation.

“What must survive compaction?” Exact commitments, approvals, operation state, unresolved work, and references to evidence; authoritative state remains outside the summary.

“Does a cached million-token prompt equal a small prompt?” No. Measure actual latency and billing for that model and workload.

Close the page and explain the difference between context, memory, and durable execution using the refund example.

A context budget with a real tradeoff

For a hypothetical 32,000-token combined budget, reserve 4,000 for output, 2,000 as headroom, 3,000 for instructions/tool schemas, 5,000 for recent dialogue, and 2,000 for selected memory. That leaves 32,000 − 4,000 − 2,000 − 3,000 − 5,000 − 2,000 = 16,000 tokens for evidence. Verify the target model's actual input/output rules; some interfaces have additional limits or reasoning-token accounting.

Suppose ten 2,000-token passages compete for that 16,000. Blindly retaining the first eight may remove the only policy exception. Rank for coverage of needed facts, preserve source/version links, remove duplicated evidence, and replace low-value text with a source-linked summary only when its qualifications survive. If all ten are essential, split the work into validated subproblems or ask a narrower question rather than pretending the budget fits.

After assembling the actual chat template, count again. Track which passage was omitted and why so a failed answer can be diagnosed. A useful memory aid is reserve first, select second, count again, verify meaning. The runnable token-budget exercise makes the accounting executable.

Interview questions with developed answers

Q1: When would you choose long context over RAG?

Sample answer: I would consider direct context when the relevant, permitted document set is bounded and the task benefits from comparing it as a whole. I would measure quality, useful recall, latency, and cost with realistic lengths and distractors. RAG is useful when the corpus is larger, changes frequently, or requires selective access, but its retrieval stage can miss evidence. A hybrid may keep a stable core in context and retrieve additional material. The choice follows the task and operating constraints; advertised context capacity alone does not settle it.

Follow-up: What if all documents fit? They may still include irrelevant or unauthorized material and produce unacceptable latency.

Q2: How do you handle high time to first token with very large prompts?

Sample answer: I measure queueing, input processing, cache behavior, and provider latency separately before optimizing. I can select less irrelevant material, reuse eligible stable prefixes, fetch details only when needed, or move suitable work to an asynchronous flow. Each change must preserve the evidence required for the task. Streaming improves perceived progress after generation starts, but does not eliminate the time before that first token. I would benchmark the actual provider and workload rather than assuming a cache hit makes a million-token prompt equivalent to a small one.

Follow-up: Can aggressive compression hurt? Yes; losing a qualification or instruction can make a faster answer wrong.

Q3: An agent works on short tasks but degrades on long ones. How do you fix it?

Sample answer: I compare traces to find lost constraints, repeated tool output, stale summaries, conflicting instructions, and missing evidence. I preserve task state and commitments explicitly, keep references to detailed sources, and compact only with a defined retention policy. I test whether the agent can recover the necessary facts after compaction and whether uncertain actions remain uncertain. I also inspect tool interfaces and no-progress loops, because context size is not the only cause. I change one policy at a time and measure completion, severe errors, calls, and cost.

Follow-up: What should never become a guessed summary? Exact approvals, payment status, and other authoritative action records.

Q4: How are context, memory, and durable execution different?

Sample answer: Context is what the model receives for one call. Memory is stored information that can be retrieved for later calls. Durable execution tracks progress and recovery across worker failures, including waits and recorded results. A communication preference belongs in memory and may be loaded into context; an approved refund and its receipt belong in authoritative workflow or business state. Mixing them lets a lossy summary overwrite facts about what the user authorized or what already happened. I connect them with explicit references and verification.

Follow-up: Does saving every chat message provide durable execution? It does not, by itself, define safe action recovery or replay.

Q5: What would you test after adding compaction?

Sample answer: I would test retention of user constraints, changes of mind, unresolved questions, source references, approvals, and completed or unknown actions. I include long histories with contradictory and adversarial content, then compare task outcomes before and after compaction. I inspect whether the summary introduces unsupported claims or upgrades untrusted text into governing instructions. I also verify that original evidence remains accessible when needed. Token savings matter only if the resulting decisions remain acceptable under the task's requirements.

Follow-up: How do you debug a bad summary? Preserve its version and provenance and compare the specific omitted or altered facts.

60-second interview answer

Context engineering is deciding what the model sees on each call: instructions, relevant history, evidence, tools, and current task state. I budget these inputs alongside the model's output and reasoning limits. I retrieve or load details when needed, summarize carefully, and keep authoritative records outside lossy summaries. Longer context and caching can help, but neither guarantees recall, correctness, or low latency. I evaluate context choices using real tasks, including long sessions, conflicting evidence, and permission changes.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Tree of Thoughts and Deliberate Search
NEXT LESSONStructured Generation →

Explore the diagram