Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

State Management Patterns

By Anup Rai7 min readReviewed September 2026

Application state is the information needed to describe a system's current condition and determine its next valid actions. For an agent task, that includes its goal, stage, artifact versions, completed operations, pending decisions and limits. It is not the model's hidden “mind,” nor necessarily the authority for external business facts.

Remember: State records where the task stands. Transition rules determine what may happen next. Persistence determines which records survive failure.

Model a task before choosing a framework

Example: an assistant prepares a database migration proposal, validates it and waits for an authorized decision before submitting it to a deployment service.

Functional requirements

  1. Track the request, target environment and proposal revision.
  2. Run independent compatibility and cost checks.
  3. Obtain approval for the validated proposal when required.
  4. Submit the permitted operation with a stable identity.
  5. Resume after a restart and report known or uncertain outcomes.

Non-functional requirements

  1. Reject invalid transitions and cross-tenant access.
  2. Handle concurrent updates without losing accepted work.
  3. Bound retained state and history size.
  4. Preserve audit evidence appropriate to the task.
  5. Support schema/code upgrades for unfinished runs.

Start with one task record and explicit transition functions. A short controller loop is reasonable. Add a graph or durable runtime when its persistence, waiting, visualization or concurrency features justify the dependency; loops are not inherently opaque and graphs are not automatically correct.

Separate state by responsibility

State class Example Update rule
Immutable task identity Tenant, run ID, original request reference Established through authenticated creation
Current workflow state Stage, proposal version, deadline Validated transitions with concurrency control
Collected results Check result keyed by task/revision Merge only compatible, deduplicated results
Conversation context Selected messages and summary Protocol-safe pruning and compaction
External operation record Submission ID and known outcome Reconcile with the owning service
Long-term memory Cross-session preference Separate lifecycle and permissions

Store large artifacts outside the frequently updated state object; keep immutable/versioned references and content hashes where appropriate. A reference must still be accessible to authorized recovery workers.

Represent valid transitions explicitly

Architecture / visual model
stateDiagram-v2 [*] --> Draft Draft --> Checking: proposal revision created Checking --> Draft: correction required Checking --> AwaitingApproval: required checks passed AwaitingApproval --> Ready: valid approval recorded AwaitingApproval --> Cancelled: rejected or expired AwaitingApproval --> Draft: proposal changed Ready --> Draft: proposal changed Ready --> Submitting: current permission rechecked Ready --> Cancelled: revoked or cancelled Submitting --> Completed: authoritative success Submitting --> Failed: confirmed rejection Submitting --> Reconciling: uncertain outcome Reconciling --> Completed: success established Reconciling --> Failed: failure established Reconciling --> Reconciling: still unknown under bounded policy Completed --> [*] Cancelled --> [*] Failed --> [*]
Read diagram source
stateDiagram-v2
    [*] --> Draft
    Draft --> Checking: proposal revision created
    Checking --> Draft: correction required
    Checking --> AwaitingApproval: required checks passed
    AwaitingApproval --> Ready: valid approval recorded
    AwaitingApproval --> Cancelled: rejected or expired
    AwaitingApproval --> Draft: proposal changed
    Ready --> Draft: proposal changed
    Ready --> Submitting: current permission rechecked
    Ready --> Cancelled: revoked or cancelled
    Submitting --> Completed: authoritative success
    Submitting --> Failed: confirmed rejection
    Submitting --> Reconciling: uncertain outcome
    Reconciling --> Completed: success established
    Reconciling --> Failed: failure established
    Reconciling --> Reconciling: still unknown under bounded policy
    Completed --> [*]
    Cancelled --> [*]
    Failed --> [*]

The approval must reference the same proposal revision that passed the checks. A changed proposal returns to validation and invalidates prior approval. An expired lease must not let an old worker submit after a new worker takes over; use the destination's supported fencing/idempotency contract.

A DAG cannot contain cycles. It suits a fixed acyclic dependency structure. A general graph/state machine can represent retries and revision cycles. Either can be orchestrated well or poorly; the distinction is structure, not a universal industry migration from one to the other.

Type the record and validate changes

Conceptual state fields:

run_id, tenant_id, schema_version, state_version
stage, proposal_ref, proposal_revision
checks_by_id, approval_ref, deadline
operation_id, external_outcome
artifact_refs, last_error, remaining_budget

Static types catch some development mistakes but do not validate untrusted runtime data by themselves. Validate schema, scope, allowed transitions and domain invariants at the update boundary. A TypedDict annotation is not an authorization check or runtime parser.

An append-only event history can support audits and reconstruction. The current materialized state can still be updated transactionally. Making every field an ever-growing list increases storage and does not by itself prevent conflicting updates.

Prevent concurrent lost updates

Two workers read version 8. One adds a compatibility result; the other adds a cost result. If both replace the whole record, the later write can erase the earlier result.

Use a transaction, per-run serialization or optimistic concurrency. Conceptual SQL:

UPDATE task_runs
SET state_json = :validated_new_state,
    state_version = state_version + 1
WHERE tenant_id = :authenticated_tenant
  AND run_id = :run_id
  AND state_version = :expected_version;

At most one update can change a given expected version under the database's applicable transaction/locking semantics. Check the affected-row count. If it is zero, reload, revalidate and recompute the state change. Do not blindly repeat an external side effect while retrying the local state update.

For parallel checks, workers can write immutable result rows keyed by (run_id, proposal_revision, check_id). The join reads results for the expected revision and checks required completion. A unique key makes duplicate delivery observable; choose whether an identical retry is ignored and whether a conflicting result is rejected or retained separately.

Design reducers and joins deliberately

LangGraph's graph API supports state channels and reducers for combining updates. Reducer choice is part of application semantics, not merely a framework convenience.

Update Possible merge Caution
Unique completed check IDs Set union Store result/version elsewhere too
Results keyed by task ID Keyed merge with conflict detection Last-write-wins can hide disagreement
Ordered conversation messages Provider-aware message reducer List concatenation can duplicate delivery
Remaining budget Atomic reservation/settlement Summing stale balances is incorrect
Approved proposal Single validated transition Approval cannot be combined across revisions

Associative/commutative reducers make some parallel combinations independent of grouping/order. Idempotence matters when updates may be delivered again. These are separate properties: integer addition is associative and commutative, but replaying the same charge twice still double-counts it.

A join must define whether it waits for all tasks, a quorum, a deadline or a minimum evidence set. “All three agents returned” is not enough if one returned an error, another checked an obsolete revision, and the third omitted required evidence.

Checkpoint and resume at real boundaries

LangGraph persistence records checkpoints at defined graph execution boundaries through a configured checkpointer. Do not assume every assignment is synchronously persisted. Backend durability, write mode and execution boundaries determine what can be lost or repeated.

On recovery:

  1. Load an authorized run and compatible state schema.
  2. Establish the latest durable progress and outstanding operation identities.
  3. Recheck current permissions and proposal validity.
  4. Reconcile uncertain external writes.
  5. Resume only the work that remains necessary and permitted.

“Resume exactly where it stopped” can be misleading: work after the last durable boundary may run again. Persist model/tool results when the old decision must be preserved. Durable execution explains the external-effect failure windows.

Treat time travel as a branch, not an undo button

Editing an earlier checkpoint and resuming can create an alternate execution path. It does not undo a database migration, refund or email already sent. Give the branch an identity, bind it to compatible artifacts and make old effect records visible to its recovery policy.

For experiments, use a read-only or simulated environment when possible. For production correction, explicitly decide which earlier outputs remain valid and which new operations are authorized. A human state edit needs the same invariant checks as an automated update.

Control state growth and upgrades

Illustrative sizing: 200 checkpoints containing a repeated 2 MB artifact produce 400 MB per run. At 10,000 runs/day, that is 4 TB/day before replication. Replacing that artifact with references and keeping each checkpoint near 8 KB reduces the same checkpoint component to 1.6 MB/run, or 16 GB/day. The external artifact still has its own storage cost, and real engines may deduplicate or use deltas.

Retain only the conversation/context needed for the next interaction, while keeping required evidence in appropriate durable records. Pruning a prompt is different from deleting the authoritative audit record.

Version state schemas and transition code. Test migrations using paused runs at different stages, including pending approvals and uncertain submissions. A deployment that can read the JSON may still interpret it incorrectly; meaning and authorization must remain compatible.

Interview practice

Q1: Is the state object the single source of truth for everything?

No. It can be authoritative for the workflow's progress, while payment/deployment systems own their business outcomes. Store references and reconcile those systems instead of treating the agent's field as proof.

Q2: Why use a graph instead of a loop?

When explicit branches, joins, waiting, persistence and inspection become easier under the graph runtime. A small loop can be simpler and fully observable. The decision depends on requirements and operational cost.

Q3: How do you merge two parallel check results safely?

Use distinct result identities and proposal revisions, apply a conflict-aware merge or immutable result rows, and let the join enforce required completion. Do not let each worker overwrite the whole shared state.

Q4: Does checkpointing prevent duplicate writes?

No. A write can succeed before the next checkpoint is saved. Stable operation identities, receiver deduplication and reconciliation address that uncertainty.

Q5: What does a zero-row optimistic update mean?

The expected state version no longer matches, the scoped record is absent, or access conditions do not match. Handle the case under the API contract; reload permitted state rather than assuming the update succeeded.

Q6: Can an append-only history still be wrong?

Yes. It can record duplicate, unauthorized or contradictory events. Validate identity and transitions, deduplicate where needed, and define deterministic projection rules.

Q7: What would you test before releasing a new state schema?

Old paused runs, concurrent updates, duplicate deliveries, stale approvals, crash windows and permission changes. Validate both recoverable state and actual external effects, plus retention and storage behavior.

Final notes

Recall card: Identity → schema → valid transition → concurrency → durable boundary → reconciliation.

Continue with framework selection after you can describe these contracts without naming a framework.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Semantic Caching
NEXT LESSONLangChain Deep Dive →

Explore the diagram