An agent control loop repeatedly chooses an action, executes it, observes the result and decides what to do next. Reasoning can guide those choices, while the application controls execution and stopping rules. A loop is a system design; hidden model reasoning alone is not an execution engine.
Different loop patterns address different task structures. They coexist rather than forming a universal progression in which every newer approach replaces ReAct.
ReAct: reasoning connected to observations
ReAct interleaves reasoning and actions, using observations from the environment to inform subsequent steps. The original work studies language models interacting with tasks and external information; it does not establish that a fixed percentage of all production agents use the method. ReAct paper.
For an illustrative build investigation:
| Step | Application-visible event | What it establishes |
|---|---|---|
| 1 | Read failed build log | The compiler reports a missing exported symbol |
| 2 | Inspect the referenced module and imports | Whether the expected symbol exists in that revision |
| 3 | Prepare a scoped edit | A candidate repair, still unverified |
| 4 | Run the relevant check | Evidence that the specific failure is fixed or persists |
The useful feedback is the observed log, code and check result. Do not require disclosure of hidden chain-of-thought to make the system observable. Record concise decisions, tool arguments, results and evidence instead.
A bounded controller
Read diagram source
flowchart TD
I[Task and acceptance criteria] --> D[Choose next permitted step]
D --> C{Action valid and budget available?}
C -->|Yes| X[Execute with operation ID and deadline]
X --> O[Record confirmed or uncertain outcome]
O --> E{Evaluate progress}
E -->|Criteria satisfied| S[Verified completion]
E -->|New useful step| D
E -->|Need user information| U[Clarification]
E -->|No progress or limit| P[Explicit incomplete result]
C -->|No| P
The controller must distinguish an observation from an instruction. Text returned by a search or a tool is evidence for the task, not permission to change the objective or execute arbitrary commands. See agent security.
Reflection and Reflexion
An evaluator-optimizer loop generates an attempt, evaluates it and revises it. Reflexion is a particular research approach using feedback and verbal reflections retained for later attempts, without updating the model's weights as part of that loop. It was introduced in 2023. Reflexion paper.
Separate the evaluator from the reflection:
- A test or reviewer reports what failed.
- The system proposes an explanation or lesson.
- The next attempt uses that lesson as a hypothesis.
- A new check determines whether the revision helped.
If a test times out, “the implementation is correct; increase the timeout” is only one hypothesis. The code may have an infinite loop, the environment may be slow, or the test may be waiting on an unavailable service. A plausible reflection should not become permanent trusted memory without supporting evidence.
Additional critic calls can improve measured outcomes, but can also reinforce shared model errors or reject correct work. Keep the acceptance criterion independent when possible, and track correction success and false rejections.
Planning before execution
Plan-and-Solve prompting first asks for a plan and then solves the problem using it. The original method is a prompting approach for reasoning tasks; a production planner/executor with tools, durable state and replanning adds further system design. Plan-and-Solve paper.
For a report comparing two service versions, a plan might load each version's measurements independently, verify workload comparability, calculate differences and prepare the report. The first two lookups can run concurrently. The comparison cannot run correctly before their results are validated.
Plans should include dependencies, required evidence and acceptance conditions. An initial plan is revisable when observations contradict it. Failure in one step need not force restarting successful independent work; replan the affected part unless dependencies or assumptions require a broader restart.
| Pattern | Useful when | Added cost or limitation |
|---|---|---|
| Reactive action/observation loop | Later choices depend heavily on new observations | Local choices may miss the broader task |
| Plan then execute | Major subtasks and dependencies can be identified early | An incorrect plan can steer all later work |
| Evaluator/optimizer | A useful quality signal can guide revision | Extra calls and evaluator mistakes |
| Explicit state graph | Allowed transitions and recovery paths need control | More application logic and state management |
These patterns can be combined. A fixed graph can contain model-directed searches and a local repair loop.
Graph orchestration does not imply determinism
A state graph names steps and allowed transitions. Nodes can run models, tools or deterministic code. The runtime can support checkpoints, interrupts and resumption; model outputs and remote side effects can still vary. LangGraph is one implementation option, not the definition of an agent. LangGraph overview.
Specify which state is persisted, how versions are handled and whether resumed execution replays any action. A graph diagram with a retry arrow is incomplete if it does not define duplicate side effects. See durable execution.
Test-time compute and external loops
Test-time compute is computation spent while producing an answer rather than updating model parameters during training. It can include longer reasoning, generating several candidates, using verifiers or searching alternatives. An external agent loop can also spend more compute and tool calls on a task. These are related resource choices, not identical mechanisms. Test-time scaling research.
Do not infer that a proprietary reasoning model internally runs Monte Carlo tree search unless its implementation is documented. Longer reasoning does not necessarily reduce external calls, latency or total cost. Evaluate the whole outcome and resource use.
Assume four attempts each consume 600 input and 200 output tokens, followed by one evaluator call consuming 900 input and 100 output tokens. Total usage is 3,300 input and 900 output tokens, before any other planning or tool-result processing. Reporting only the final 200-token answer would conceal most of the work.
Handle the failures a loop creates
| Failure | Detection | Repair and tradeoff |
|---|---|---|
| Repeated query returns identical evidence | Track normalized query, parameters and result IDs | Stop or change a justified constraint; avoid blocking useful retries after a transient outage |
| Plan relies on a false premise | Compare assumptions with observations | Replan affected dependencies |
| Tool timeout leaves outcome unknown | Preserve operation ID and status | Reconcile before retrying a consequential action |
| Critic rewards a wrong result | Compare with independent checks and reviewed labels | Improve rubric or evaluator; more review cost |
| Context compaction loses constraints | Validate structured state after compaction | Restore constraints/evidence; extra state management |
| Parallel calls conflict | Identify shared writes and dependency boundaries | Serialize, lock or version-check those operations |
“Never retry the same tool” is too broad: transient failures may justify bounded retry. Conversely, retrying a non-idempotent action without reconciliation can duplicate effects. Classify the failure before choosing a recovery action.
Evaluate loop quality
Measure task success, critical violations, unnecessary actions, repeated failures, recovery quality, total tokens and wall-clock latency. Compare with a simpler workflow on the same cases. Include missing information, contradictory observations, external outages and resumed tasks.
Set a maximum number of actions and a no-progress rule, but do not force the model to claim success when either limit is reached. Return an incomplete outcome with the evidence collected and remaining requirement. Where a user decision is necessary, ask a targeted question and retain recoverable state.
Interview practice
Q1: When would you use ReAct instead of an initial plan?
When useful next actions depend strongly on observations that are not available upfront. I may still maintain a high-level plan and acceptance criteria. Reactive execution and planning are compatible; the choice is how much structure can be determined reliably in advance.
Q2: Does Reflexion train the model during the task?
The studied loop stores verbal feedback for later attempts rather than updating model weights. A generated lesson can be wrong, so I preserve its source and validate whether it improves subsequent attempts.
Q3: Why does a state graph not make the agent deterministic?
The graph constrains transitions, but its nodes may call stochastic models and changing external services. Checkpoints help recovery; they do not automatically make retries safe or remote effects exactly once.
Q4: How much should the agent replan after failure?
Enough to repair invalid assumptions and affected dependencies. Preserve successful independent work when still valid. A full restart can waste resources or repeat side effects, while a narrow repair can fail if the original premise is wrong.
Q5: Does more reasoning mean fewer tool calls?
Sometimes, but it is an empirical outcome rather than a guarantee. Extra internal reasoning can also make a wrong plan more elaborate. Measure task success, latency and total cost with the selected model and tool environment.
Q6: What belongs in a trace?
Task/version identifiers, relevant state changes, proposed and executed actions, operation IDs, observations, timing, budget use and verified outcomes. Concise rationale can help debugging; hidden reasoning is not required to establish which actions occurred.
Q7: How should the loop stop?
On verified completion, required clarification, cancellation, exhausted limits or lack of useful progress. Distinguish those statuses so downstream systems and users do not mistake an incomplete run for success.
Final notes
Recall card: Choose a step → execute under constraints → inspect the observation → update state → verify or continue. A good loop makes progress testable and failure recoverable.