Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Error Handling and Recovery

By Anup Rai10 min readReviewed September 2026

Numerical examples are illustrative unless explicitly sourced.

Error handling detects and classifies failures, then selects an appropriate response. Recovery restores the task to a known valid state or ends it with an accurate account of what remains unresolved. A successful retry is only one possible recovery; stopping a denied action or reconciling an uncertain write can be the correct result.

Remember: Classify → contain → recover → verify.

Learn to classify a failure before choosing a recovery

An agent is software, so ordinary exception handling remains necessary. What changes is that successful execution of a function is not always successful completion of the user's task. A search tool can return HTTP 200 with no useful evidence; a payment call can time out after charging; a model can repeatedly call a real tool with the wrong argument.

Separate four questions. Did the request reach the service? Did the operation complete? Was the result valid for the task? Did the overall task achieve its intended outcome? Different evidence answers each question. A try/catch block can catch a network exception, but it cannot prove that the returned policy applies to this customer.

Classify errors into categories with explicit responses: transient transport failure, invalid input or schema, authorization denial, invalid business result, uncertain side effect, and lack of progress. Some can be retried, some need corrected input, some must stop, and some require external reconciliation. Sending every failure back to the model with “try harder” ignores those distinctions.

Follow a failed booking attempt

The agent proposes a date in the wrong format. The executor rejects the schema before making a reservation and returns a concise typed error showing the accepted format. The agent can repair the argument within a bounded correction budget. No reservation was attempted, so this is different from a timeout after submission.

On the next attempt, the booking service reports that the room is unavailable. This is a valid business outcome, not a broken API. The agent may search alternatives within the user's constraints or ask for a changed requirement. It should not repeat the same unavailable booking indefinitely.

Now suppose the reservation request times out. It may have succeeded. Preserve the same operation identity and query the authoritative reservation state or use the receiver's deduplicating retry contract. Do not create a new reservation just because the receipt was not received. If the outcome remains unknown, pause for reconciliation.

A useful tool result separates status and evidence:

status: rejected | completed | unknown
error_class: invalid_argument | unavailable | permission_denied | transient
operation_id: stable identity when a write may have been attempted
retry_allowed: policy decision with conditions, not model permission
evidence: receipt or authoritative lookup reference when available

This is a conceptual contract; actual implementations should expose only the fields and data the caller is allowed to see.

Recover the plan without losing control

Model-assisted correction can use an actionable error to choose a different permitted plan. Give it enough information to fix the problem without leaking credentials or treating untrusted error text as governing instructions. Enforce attempt, time, and spending budgets outside the model.

Detect loops by observing repeated actions and lack of meaningful state change, not only identical strings. Two differently worded searches can repeat the same failed strategy. Conversely, a repeated status query can be legitimate polling. Define progress in terms of the task and use a bounded wait or handoff when it stalls.

Checkpoints preserve recorded state for resumption. They do not reverse real-world effects. A recovery plan should say which work can repeat, which results can be reused, and which outcomes need verification. Remember classify, contain, correct, verify: the model can help correct a plan, while the application remains responsible for containment and verification.

One failure, several possible meanings

A tool says “request failed.” Before asking the model to try harder, ask what happened. An invalid order ID is different from an unavailable server, and neither establishes whether a payment committed.

Failure First response Avoid
Invalid schema or missing field Validate; allow a limited repair with a clear error Repeatedly sending the same malformed request
Authentication/authorization denied Stop and resolve identity or permission through the application Asking the agent to find a bypass
Rate limit or temporary outage Backoff with jitter within a shared retry/deadline budget Nested retries multiplying load
Read returned incomplete or stale data Check version, completeness, and authoritative source Equating HTTP 200 with correct data
Write timed out Mark outcome unknown and reconcile Retrying with a new operation ID
Valid tool calls but no progress Detect repeats; stop, change strategy, or hand off An unlimited self-correction loop

Make the recovery decision explicit

Architecture / visual model
flowchart TD F[Failed or questionable outcome] --> C[Classify using execution evidence] C -->|Invalid arguments| A[Bounded correction and validation] C -->|Temporary safe-to-retry failure| R[Deadline-aware backoff] C -->|Denied operation| S[Stop action and report required authority] C -->|Uncertain external write| Q[Reconcile stable operation ID] C -->|Valid output but no progress| P[Revise strategy within budget] A --> V[Verify resulting business outcome] R --> V Q --> V P --> V V -->|Known complete| D[Record completion] V -->|Still unresolved| H[Return partial result or operator handoff]
Read diagram source
flowchart TD
    F[Failed or questionable outcome] --> C[Classify using execution evidence]
    C -->|Invalid arguments| A[Bounded correction and validation]
    C -->|Temporary safe-to-retry failure| R[Deadline-aware backoff]
    C -->|Denied operation| S[Stop action and report required authority]
    C -->|Uncertain external write| Q[Reconcile stable operation ID]
    C -->|Valid output but no progress| P[Revise strategy within budget]
    A --> V[Verify resulting business outcome]
    R --> V
    Q --> V
    P --> V
    V -->|Known complete| D[Record completion]
    V -->|Still unresolved| H[Return partial result or operator handoff]

Store job ID, input version, completed results, operation IDs, pending approvals, attempt count, deadline and error category in a durable backend if recovery must survive a restart. The sequence is:

  1. Preserve the last known state and any possible side effect.
  2. Choose recovery from the error semantics.
  3. Check remaining authority, time and resource budget.
  4. Perform the bounded recovery step.
  5. Verify and record its outcome, including uncertainty.

Prevent retries from amplifying an outage

Choose one layer to own retries where practical. If four nested layers each allow three attempts, one top-level request can trigger 3⁴ = 81 calls to the failing dependency. “Three attempts” includes the original call; “three retries” would mean four attempts per layer.

For an eligible read, an example exponential-backoff policy samples a delay uniformly between zero and min(cap, base × 2^retryIndex). With a 200 ms base, the first three upper bounds are 200, 400 and 800 ms. These are chosen example values, not a standard for every service. Honor applicable server retry guidance and stop when the remaining deadline cannot accommodate another attempt. Check SDK defaults so hidden retries do not exceed the shared budget.

A circuit breaker can temporarily stop calls to an unhealthy dependency; a concurrency limit contains in-flight work; a retry budget limits extra load. These address different problems. Scope controls to the relevant dependency or tenant, and define recovery probes so one failing route does not unnecessarily disable unrelated work. A fallback must preserve the required semantics: cached availability is not proof that a room can still be booked.

Retries of writes require an established idempotency or reconciliation contract, regardless of the backoff policy. See Temporal activity retry and idempotency semantics.

What self-correction can and cannot do

A model can propose a better search query or correct a malformed argument. Give it the relevant, sanitized error and remaining budget. Do not feed arbitrary raw exception text back as trusted instructions: errors can contain secrets or attacker-controlled content.

A verifier model is an additional signal. For a financial total, recompute arithmetic; for a reservation, query the booking system; for an access check, call the authorization service. A fluent explanation is not verification.

Checkpoint restoration rewinds application state. It does not undo emails, payments, or database writes. A compensating action is a separate operation with its own authorization, idempotency, and failure handling. Temporal activity semantics.

Prevent the endless loop

Use several independent limits: wall-clock deadline, maximum attempts per operation, total tool calls, token/cost budget, and lack-of-progress detection. Count equivalent normalized requests, not only identical strings. A changed search spelling can still be the same failed strategy.

A budget should fit the task. Ten steps is not a universal limit for a two-minute support task and a multi-hour research job. Bound parallel work too: stopping the parent must propagate cancellation where possible, while already-committed effects still need reconciliation.

The manager decision

For a low-risk answer, return a partial result with clear missing information. For an uncertain irreversible action, pause. Staff the exception queue and give the operator the intended action, actual evidence, attempted operations, known effects, and safe next choices. Do not hand over only a huge chat transcript.

Track first-attempt success, recovery success, unresolved outcomes, duplicate effects, manual minutes, and total cost per successful task. A high recovery rate can conceal a broken upstream dependency; fix repeated root causes rather than celebrating retries.

Practice without looking

A verifier approves a wrong total—what changes? Add deterministic arithmetic and source reconciliation; evaluate the verifier against labeled mistakes.

The worker restarts after sending an email—what now? Inspect the provider's delivery/operation record and deduplication contract. A local checkpoint cannot unsend it.

The agent keeps trying a denied tool—what stops it? The tool gateway rejects it regardless of model output, and the orchestrator ends the attempt under policy.

See Durable execution for the full crash-window walkthrough and LangGraph persistence for persistence boundaries.

Interview questions with developed answers

Q1: Why is try/catch alone insufficient for agent recovery?

Sample answer: It handles program exceptions, which remain important, but many task failures are valid-looking results or uncertain business outcomes. An empty search response or an incorrect invoice total may not raise an exception. I define typed error and outcome categories, validate results, and choose recovery by semantics. The model can repair some arguments or revise a plan, but permission denials stop the action and ambiguous writes require reconciliation. Recovery stays bounded by time, attempts, and cost. Exception handling and agent-level recovery complement each other rather than replacing one another.

Follow-up: Which errors should never simply be retried unchanged? Invalid inputs, authorization denials, and uncertain writes without a safe receiver contract.

Q2: How do you handle silent failures where a tool returns 200 but the result is wrong?

Sample answer: I validate the tool's business result, not just its HTTP status. For structured data I check schema, identifiers, ranges, and domain invariants. For a write I verify authoritative state or a receipt. For semantic results I use evidence checks or calibrated review as appropriate. A verifier model can help with some judgments but is also fallible, so it does not replace deterministic checks where those are available. I record the failure category and add a regression case that checks the actual outcome.

Follow-up: How would you validate an invoice total? Recompute it from verified line items, taxes, and discounts under the applicable rules.

Q3: How do you stop an agent from looping?

Sample answer: I enforce step, time, token, and tool-call budgets in the orchestrator and monitor whether the task state is advancing. I inspect repeated actions and similar failed strategies, then provide a bounded opportunity to change the plan when a permitted alternative exists. If no progress is possible, I stop with an honest partial result or handoff that includes completed work and blockers. I also fix poor tool feedback or ambiguous objectives causing the loop. A prompt saying “do not loop” is useful guidance but not a resource limit.

Follow-up: Can repetition be valid? Yes, for bounded polling or retries under a defined policy.

Q4: What should happen after a payment timeout?

Sample answer: I mark the outcome unknown and preserve the intended operation ID. I query the payment service or retry with the same key only under its documented deduplication contract. If it reports completion, I recover the receipt and continue. If the result cannot be established, I pause for reconciliation instead of issuing a new payment. I also ensure concurrent workers cannot start independent duplicates. The user's message should reflect the uncertainty rather than falsely claiming either success or failure.

Follow-up: Does cancelling the task prove no payment occurred? No; the remote effect may already have committed.

Q5: What information makes a human handoff effective?

Sample answer: The handoff should state the user goal, constraints, completed and uncertain actions, evidence references, attempted recovery, and the exact decision or authority needed next. It should preserve operation IDs so the reviewer does not repeat an action unknowingly. Sensitive data is shared only with an authorized reviewer. I also provide a deadline or priority and an owner for the queue. A generic “agent failed” message forces the human to reconstruct the task and increases both delay and duplicate-action risk.

Follow-up: How do you measure handoff quality? Resolution time, avoidable rework, duplicate actions, and reviewer feedback on missing information.

60-second interview answer

I separate transient failures, invalid requests, permission failures, incorrect outputs, and unknown external outcomes. Each needs a different response. A temporary read failure may deserve a bounded retry; a denied action should stop; a timed-out payment needs reconciliation. The model may help repair a plan, but code enforces permissions, budgets, and business invariants. I persist useful progress, stop repeated unproductive attempts, and provide a human handoff with the last known state. Recovery leaves a verified result or a clearly recorded unresolved state; it never invents success.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Planning and Decomposition
NEXT LESSONHuman-in-the-Loop Patterns →

Explore the diagram