Planning selects actions and their ordering to reach a goal under constraints. Decomposition breaks a task into smaller tasks with explicit dependencies. A plan can be a short checklist, a dependency graph or a policy for choosing the next action as observations arrive.
A model-generated plan is a proposal. The runtime must still check feasibility, permissions, deadlines and outcomes. Describing planning as “System 2” is an analogy; it does not establish how a particular model works internally or guarantee logical correctness.
Choose the simplest useful planning method
| Method | How it works | Suitable situation | Limitation |
|---|---|---|---|
| Fixed workflow | Follow authored steps and conditions | Stable process with known exceptions | New situations need explicit handling |
| Plan then execute | Produce an initial sequence or dependency graph | Goal is clear and dependencies are predictable | Assumptions can become stale |
| Adaptive planning | Revise affected work after new evidence | Environment or requirements change | Extra calls and plan churn |
| Hierarchical decomposition | Split a goal into bounded subgoals | Several distinct responsibilities | Missing shared assumptions and coordination cost |
| Search over alternatives | Generate, evaluate and compare candidate paths | Reliable evaluation and affordable exploration exist | Search cost and evaluator errors |
Plan-and-Solve is a prompting method that first decomposes a reasoning task and then solves the subtasks. It does not prescribe an immutable ten-step tool workflow, automatic parallelism or a production recovery engine. Original paper.
Work from a real task and observable outcomes
Task: prepare a pull request that updates an API client after a schema change. Publishing a release is outside this example's authorized scope.
Functional requirements
- Read the current API schema and affected client code.
- Identify required compatibility changes and tests.
- Implement the agreed scope.
- Run relevant validation and report failures accurately.
- Prepare a reviewable diff and explain remaining risks.
Non-functional requirements
- Preserve unrelated owner edits.
- Bound tool/model calls and elapsed time.
- Isolate test execution and use only authorized credentials.
- Record input revisions and completed artifacts.
- Resume without repeating unsafe side effects or accepting stale results.
Baseline: inspect, edit, test and summarize sequentially. This is easy to follow. Its flaw is that “inspect” hides several independent activities, while “edit” may begin before compatibility requirements are understood.
Make dependencies explicit
Read diagram source
flowchart TD
G[Goal and scope] --> S[Read schema revision]
G --> C[Inspect client and existing tests]
S --> D[Determine contract changes]
C --> D
D --> I[Implement client changes]
D --> T[Prepare required tests]
I --> V[Run validation against current diff]
T --> V
V -->|Pass| R[Prepare pull request summary]
V -->|Failure with actionable evidence| F[Revise affected work]
F --> I
V -->|Blocked or budget exhausted| B[Report exact incomplete outcome]
The initial schema and code inspection can run independently. Test preparation may overlap implementation once the contract is agreed. The validation step depends on both. Repeating a failed test without changing its cause does not advance the plan.
The dependency portion can be represented as a DAG. The complete execution controller has a recovery cycle, so the full diagram is not a DAG. Decomposition does not require a separate model or agent for every box.
Represent each task with a contract
| Field | Purpose | Example |
|---|---|---|
| ID | Stable reference across revisions | validate-client |
| Inputs and revisions | Establish what the result applies to | Schema revision and source commit |
| Dependencies | Prevent execution before prerequisites | Implementation and tests completed |
| Preconditions | Check required environment and scope | Test dependencies available |
| Action | Bounded executable work | Run the selected test suite |
| Success evidence | Observable completion criterion | Exit result and relevant test report |
| Side effects | Define safety/retry behavior | Writes temporary build artifacts |
| Deadline/budget | Bound execution | Remaining task time and calls |
| Status and artifact | Support resume and audit | Failed, with report ID |
“Tests look good” is weaker than a record of which tests ran against which diff. A successful old test report does not establish that a later code revision passes.
Before executing a generated plan, check that required inputs exist, dependencies are satisfiable, tools are available and actions remain within scope. Reject self-dependencies, missing task IDs and contradictory acceptance criteria. Model planning and deterministic validation complement each other.
Replan the affected region
Suppose schema inspection succeeds, but the client test reveals a field changed from optional to required. Preserve the schema evidence and update the implementation and dependent tests. Do not discard unrelated successful investigation merely because one step failed.
- Record the observed failure and current input revisions.
- Distinguish a temporary execution problem from an invalid plan assumption.
- Find tasks whose inputs or outputs are affected.
- Reuse still-valid completed artifacts.
- Revise the remaining plan and budget.
- Verify the final result against the revised current state.
Checkpointing supports recovery of execution state. It does not undo an email, credit or deployment already performed by an external system. Such operations need their own idempotency, reconciliation or compensation contract; see durable execution.
Plan revision is not inherently more expensive than initial planning. A targeted correction can be cheap; rebuilding a large plan from a verbose transcript can be expensive. Store compact task records and dependencies so revision does not require rediscovering all completed work.
Bound decomposition and search
Depth alone does not control expansion. With three children per task and three levels below the root, a fully expanded tree contains 1 + 3 + 9 + 27 = 40 tasks. At five children per task, the same depth permits 156 tasks.
Control total task count, fan-out, concurrent workers, token/tool budget and deadline. Each child receives a bounded allocation from the parent. Stop decomposing when a task already has a clear executable action and success criterion; do not ask a model to invent subtasks merely to fill a hierarchy.
| Control | What it prevents | Remaining limit |
|---|---|---|
| Maximum depth | Unbounded nesting | Wide shallow trees can still explode |
| Total task budget | Excessive total work | Tasks can differ greatly in cost |
| Concurrency limit | Resource bursts | Queuing can miss deadlines |
| No-progress detector | Repeated ineffective actions | Legitimate slow progress needs suitable evidence |
| Outcome-based stop rule | Continuing after success | Success verification can be wrong |
Minimal worker context should mean sufficient and relevant, not merely short. A worker needs the task objective, constraints, input versions and expected output. Omitting an essential dependency increases error risk even if it saves tokens.
Understand tree search before proposing MCTS
Tree of Thoughts explores alternative reasoning states using an explicit search procedure. It is not necessarily Monte Carlo Tree Search. MCTS repeatedly performs selection, expansion, evaluation/rollout and backup of value estimates, balancing exploration with exploitation. Simply asking for ten suggestions and choosing the highest model score is candidate ranking, not a full MCTS implementation. Tree of Thoughts.
Language Agent Tree Search integrates MCTS with model proposals, value estimates, reflection and environment feedback. It provides evidence for a particular research method, not proof that all commercial reasoning models internally run MCTS. LATS.
| Search ingredient | Example | Question to answer |
|---|---|---|
| State | Candidate code patch and known test results | Can it be copied/restored safely? |
| Action | A proposed patch variation | Is it valid and within scope? |
| Transition | Apply patch in an isolated workspace | Is the environment faithful? |
| Evaluation | Tests plus reviewed requirements | Can the candidate exploit the evaluator? |
| Budget | At most a bounded number of evaluations | Is expected improvement worth the cost? |
An LLM imagining a tool result is not equivalent to observing the real environment. Exploration is most defensible when candidates can be evaluated safely in an isolated or simulated environment. Do not “try several alternatives” by issuing real payments or publishing multiple external changes.
Search can amplify evaluator mistakes: more candidates may provide more opportunities to find a misleading high score. Keep held-out checks and inspect whether the selected plan satisfies the original goal. High consequence alone does not make search appropriate; evaluator quality and safe experimentation matter.
Interview practice
Q1: How do planning and reasoning differ?
Reasoning may help infer facts or evaluate alternatives; planning specifies actions toward a goal. An application should expose an actionable plan and completion evidence without depending on hidden model reasoning as its durable state.
Q2: When would you use a fixed workflow?
When steps and important exceptions are known and the process benefits from predictable control. Add model decisions only where interpreting inputs or choosing among legitimate alternatives needs them. A workflow can include branches and retries.
Q3: Does decomposition create parallelism?
Only if the subtasks are actually independent. Represent data dependencies and shared write constraints. Two tasks with different names can still depend on the same changing resource.
Q4: What happens when step two fails?
Classify the failure, preserve still-valid artifacts and replan affected dependent work. A timeout may need bounded retry; a changed assumption may need a new plan. A checkpoint cannot reverse an already completed external action.
Q5: How do you prevent recursive expansion?
Use a global budget plus depth, fan-out and concurrency limits. Validate task granularity and require observable progress. Depth three alone can still create many tasks.
Q6: When would MCTS be reasonable?
When there is a meaningful state/action model, a useful evaluator, safe exploration and enough budget. Explain selection, expansion, evaluation and backup. Compare it with cheaper candidate generation or a single adaptive plan.
Q7: How do you evaluate a planner?
Measure final task success, invalid dependencies, missing requirements, unnecessary work, replanning frequency, cost and latency. Separate a bad plan from a correct plan whose execution failed. Test changing inputs and partial completion, not only clean first attempts.
Final notes
Recall card: Goal → constraints → dependencies → executable tasks → observed evidence → targeted revision. Close the interview with the chosen planning method, its budget, recovery behavior and how you will establish completion.