LangGraph is an orchestration framework for stateful workflows and agents. It represents work through state, nodes and control flow, with facilities for persistence, interrupts and streaming. A graph can contain ordinary deterministic code, model calls or both; using LangGraph does not make every workflow an autonomous agent. LangGraph overview.
Use it when explicit transitions, resumable work or branching make the application easier to reason about. Adoption statistics and GitHub stars are not evidence that it fits a particular reliability requirement.
Learn the execution vocabulary
| Term | Meaning | Interview example |
|---|---|---|
| State | Values carried through an execution | Proposal revision and validation results |
| Node | Function performing a unit of work | Validate a proposed configuration |
| Edge | Control-flow connection | Validate before deciding the next step |
| Conditional routing | Select next work from current information | Request review or finish |
| Reducer | Rule combining updates to a state field | Merge distinct check results |
| Checkpoint | Stored graph state at an execution boundary | Recover a paused proposal |
| Thread ID | Identifier grouping persistent graph history | One authorized task/conversation |
| Interrupt | Pause for external input | Wait for a decision on a specific proposal |
LangGraph supports cycles; it does not require one. Current LangChain agents also use LangGraph, so “LangChain is always acyclic, LangGraph is cyclic” is an incorrect comparison. A DAG is an acyclic structure; a retry/revision loop requires a cycle or another explicit repetition mechanism.
Start from a deployment-proposal workflow
Functional requirements
- Prepare a configuration proposal against a specified environment revision.
- Run schema and operational checks.
- Revise failed proposals within a bounded attempt budget.
- Obtain any required human decision tied to the exact proposal.
- Submit through an authorized deployment service and report the actual outcome.
Non-functional requirements
- Enforce tenant/environment access at execution and resume.
- Persist required progress across worker failure.
- Avoid duplicate submission after ambiguous responses.
- Bound retries, elapsed time and model/tool usage.
- Keep unfinished runs compatible with deployments of new code.
Read diagram source
flowchart TD
A[Authenticated goal and environment revision] --> P[Prepare proposal]
P --> V[Validate proposal and required checks]
V --> D{Valid?}
D -->|No, budget remains| P
D -->|No, exhausted| F[Record failure with evidence]
D -->|Yes| H{New approval required?}
H -->|Yes| I[Persist proposal and interrupt]
I -->|Accepted decision| G[Recheck proposal, permission and expiry]
I -->|Rejected or expired| X[Stop proposed action]
H -->|No, already authorized| G
G -->|Permitted| T[Submit with stable operation ID]
G -->|Stale or denied| X
T --> O{Authoritative outcome}
O -->|Success| E[Persist receipt and finish]
O -->|Confirmed failure| F
O -->|Unknown| R[Reconcile under bounded policy]
The graph makes control flow visible. The tool service still owns authentication, authorization, idempotency and deployment status. A Boolean is_secure in graph state cannot replace those checks.
Build the smallest graph first
This provider-free example demonstrates state and conditional routing. It needs a compatible LangGraph installation and performs no external action:
from typing import TypedDict
from langgraph.graph import END, START, StateGraph
class ReviewState(TypedDict):
proposal: str
valid: bool
status: str
def validate(state: ReviewState):
return {"valid": bool(state["proposal"].strip())}
def ready(state: ReviewState):
return {"status": "ready_for_review"}
def reject(state: ReviewState):
return {"status": "missing_proposal"}
builder = StateGraph(ReviewState)
builder.add_node("validate", validate)
builder.add_node("ready", ready)
builder.add_node("reject", reject)
builder.add_edge(START, "validate")
builder.add_conditional_edges(
"validate",
lambda state: "ready" if state["valid"] else "reject",
{"ready": "ready", "reject": "reject"},
)
builder.add_edge("ready", END)
builder.add_edge("reject", END)
graph = builder.compile()
result = graph.invoke({"proposal": "", "valid": False, "status": "new"})
assert result["status"] == "missing_proposal"
This is a deterministic workflow, not an agent. Real proposal validation checks schema and domain invariants; nonempty text is only the demonstration's routing condition. Model-based proposal generation would be another node with its own failure and budget contract.
Understand updates and message reducers
Nodes commonly return partial updates. The state field's reducer determines whether those updates replace or combine with existing data. Independent parallel tasks should not silently overwrite one shared field.
add_messages is more than append-only list concatenation: it recognizes message IDs and can replace an existing message with the same ID. Stored messages are observable interaction records, not guaranteed access to a model's complete internal reasoning. See state and reducer semantics.
For parallel schema and capacity checks, key each result by check ID and proposal revision. The join requires both successful results for the same revision. A generic list reducer does not establish those invariants. State-management patterns covers conflict-aware merging and optimistic updates.
Explain what persistence actually preserves
Checkpointers organize state by thread and full checkpoints at super-step boundaries. A super-step groups the nodes scheduled for one execution tick. Per-task pending writes can preserve successful peer-node outputs when another node fails in that step.
Persistence still depends on the backend and durability mode. An in-memory saver is useful for local examples; it does not survive process loss. An opaque thread ID selects history but is not a permission token: authenticate access to resume, inspect or edit it.
For illustration, schema validation completes, capacity validation fails, and the runtime retains the successful task write. Recovery may reuse schema validation instead of rerunning it. If the proposal changes, application semantics may require invalidating that result anyway. Runtime recovery and domain validity are separate questions.
Do not assume every Python assignment is durably saved. Choose persistence mode according to acceptable lost progress and latency, and test the configured backend at the relevant failure boundaries. Durable execution compares sync, async and exit persistence.
Resume an approval safely
LangGraph's interrupt() exposes a JSON-serializable request for external input. Resume uses Command(resume=...) with the same thread identity. A crucial detail: the interrupted node starts again from its beginning, so code before the interrupt can execute again. Interrupt contract.
Design the application around this behavior:
- Create and persist an identifiable proposal before requesting review.
- Keep pre-interrupt code safe to repeat.
- Authenticate the person submitting the decision.
- Bind the decision to proposal revision, scope and expiry.
- Recheck current conditions before the external action.
- Route rejection or expiration to an explicit stopped outcome.
The application supplies expiry handling; an indefinite framework wait is not an approval deadline. Approval already granted for the applicable action does not need to be requested again solely because a worker restarted.
Time travel and external effects
Time travel lets developers inspect checkpoints and explore alternate continuations. Editing old state does not reverse an earlier deployment. Work after the selected boundary can execute anew, so use simulations or appropriately idempotent side-effect boundaries when debugging.
Keep a link between the branch, proposal and earlier operation IDs. If the intended operation changes, decide whether it is a new authorized business action; do not silently reuse an old approval or generate fresh IDs for an uncertain retry.
Use subgraphs and multiple agents for a reason
| Pattern | Useful when | Additional concern |
|---|---|---|
| Supervisor with workers | Independent investigations need integration | Delegation scope, shared budget and evidence quality |
| Handoff | A different specialist should continue the task | Transfer sufficient state and preserve authority |
| Subgraph | A reusable workflow has its own local state | Schema mapping, persistence and parent/child boundaries |
| Parallel checks | Independent checks reduce elapsed time | Fan-out limits and revision-aware joining |
Limit the state and tools passed to each worker. However, a private graph channel is not automatically a security boundary: streaming, tracing, checkpoint access and runtime credentials must also be considered. A subgraph with access to broad credentials can still exceed the intended role.
Use a bounded loop condition, deadline and operation budget. A recursion/step guard is useful but does not on its own cap wall-clock time or money spent in one long-running node.
Test behavior before choosing deployment
- Validate nodes against malformed and unauthorized input.
- Test routes, loop limits and join conditions.
- Send duplicate, stale and rejected approval decisions.
- Crash during model calls, pending writes and external submissions.
- Resume using a different worker and a compatible new code version.
- Check streamed events and traces for prohibited data.
- Measure success, recovery time, tail latency and full cost.
Local execution, a self-managed service and a managed agent platform have different operating responsibilities. Using an open-source graph does not automatically make model calls, telemetry or storage stay on premises. Inspect the entire configured data path.
Interview practice
Q1: Is every LangGraph application multi-agent?
No. A graph can be a deterministic workflow or one agent with tools. Add multiple agents only when independent contexts or responsibilities justify the coordination.
Q2: What does add_messages preserve?
It merges message records using message identities, adding new messages and updating matching ones. It is not an immutable event log or proof of the model's internal reasoning process.
Q3: Does a thread ID secure a conversation?
No. It identifies persisted history. The application must authorize every read, resume, mutation and export associated with that history.
Q4: Why avoid an irreversible action before interrupt()?
The node restarts on resume, so that action may run again. Place it behind an appropriate task/effect boundary with deduplication and current authorization.
Q5: Why can replaying a checkpoint change the outside world?
Later nodes can execute again. Checkpoints restore application state, not external services. Use safe test environments and explicit side-effect semantics.
Q6: What is the strongest reason to adopt LangGraph?
It makes the required state transitions, branching, waiting and recovery easier to implement and inspect under tested contracts. Popularity alone is not a reliability argument.
Final notes
Recall card: State → transition → reducer → checkpoint → authorized resume → verified external outcome.
Next: LangSmith observability.