Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Multi-Agent Orchestration

By Anup Rai8 min readReviewed September 2026

Multi-agent orchestration coordinates two or more agents: it assigns work, controls communication, manages shared state and combines results. Each agent may have its own model, instructions, tools and context. Multiple model calls in a fixed pipeline are not necessarily multiple autonomous agents.

The interview decision is whether separate agents improve a measurable outcome enough to justify coordination cost. A single agent with well-designed tools, or a conventional workflow, is often a sufficient starting point.

Decide what needs to be separated

Consider an assistant preparing a service migration proposal. It must inspect dependencies, estimate infrastructure cost and assess operational readiness. These investigations can often proceed independently against a common service revision. The final proposal needs one owner to reconcile assumptions and contradictions.

Useful reasons for separate agents include:

  1. Independent investigations that can execute in parallel.
  2. Different tools or permissions for different responsibilities.
  3. Large evidence sets that benefit from separate contexts and compact artifacts.
  4. Different evaluation criteria, model choices or operational owners.

More agents do not inherently improve quality. If every worker needs the full conversation and edits the same artifact, communication and conflicts can outweigh the benefit. Anthropic's research-system report describes benefits for parallel investigations and substantial token overhead; those observations concern its workload, not a universal multiplier. Engineering report.

Compare orchestration patterns

Pattern Control structure Useful when Main cost or risk
Pipeline Ordered stages with explicit inputs/outputs Dependencies are known An early error propagates downstream
Supervisor and workers Coordinator delegates and integrates Work can be decomposed dynamically Poor decomposition or coordinator bottleneck
Parallel workers with reducer Independent tasks, then deterministic merge or synthesis Breadth matters and tasks are separable Stragglers, duplicated work and integration errors
Handoff Active responsibility transfers to another agent A conversation changes specialist/domain Lost context, loops or unclear ownership
Peer discussion/debate Agents exchange proposals and critiques Independent alternatives are valuable Correlated mistakes and persuasive wrong consensus
Hierarchy Supervisors delegate through several levels Scope exceeds one coordinator's manageable span Extra latency and distortion at each level

A directed graph can represent these patterns, including cycles. A DAG is specifically acyclic. Graphs do not imply deterministic model outputs, and a conditional retry in a workflow does not automatically require another agent. LangGraph explicitly models state, functions and transitions and supports loops. Graph API.

Start with a baseline, then add bounded parallelism

Baseline: one worker gathers all evidence sequentially and writes a proposal. Measure omitted dependencies, unsupported claims, completion time and cost. Split investigations only where those measurements suggest a benefit.

Architecture / visual model
flowchart TD R[Request and authorized scope] --> P[Plan and dependency check] P --> B[Reserve shared task budget] B --> D[Dependency investigation] B --> C[Cost investigation] B --> O[Operations investigation] D --> A[Versioned evidence artifacts] C --> A O --> A A --> V[Validate coverage and reconcile conflicts] V -->|Evidence sufficient| S[Produce proposal with sources] V -->|Specific gap and budget remains| F[One targeted follow-up] F --> A V -->|Blocked or budget exhausted| E[Report limits and unresolved decisions]
Read diagram source
flowchart TD
    R[Request and authorized scope] --> P[Plan and dependency check]
    P --> B[Reserve shared task budget]
    B --> D[Dependency investigation]
    B --> C[Cost investigation]
    B --> O[Operations investigation]
    D --> A[Versioned evidence artifacts]
    C --> A
    O --> A
    A --> V[Validate coverage and reconcile conflicts]
    V -->|Evidence sufficient| S[Produce proposal with sources]
    V -->|Specific gap and budget remains| F[One targeted follow-up]
    F --> A
    V -->|Blocked or budget exhausted| E[Report limits and unresolved decisions]

The workers do not all edit the final proposal. They return findings tied to the same service snapshot. The coordinator checks whether a cost estimate used a different traffic assumption or whether an operations recommendation contradicts a dependency constraint.

Functional requirements

  1. Accept a task with its scope, required outputs and success criteria.
  2. Assign independent subtasks with explicit input dependencies.
  3. Collect evidence, status and artifacts from each worker.
  4. Resolve contradictions or report them as unresolved.
  5. Support cancellation, targeted retries and partial results.

Non-functional requirements

  1. Enforce one overall deadline and budget across all descendants.
  2. Preserve tenant and object-level permissions during delegation.
  3. Make completed artifacts durable and traceable to their inputs.
  4. Bound fan-out, depth, concurrent tool calls and retry counts.
  5. Recover without duplicating consequential actions.

These are design requirements, not promises supplied by a framework name.

Define a delegation contract

Field Example for the cost worker
Task ID and parent cost-17, parent migration-8
Objective Estimate cost for the proposed traffic envelope
Inputs Service revision, volume assumptions, region and approved price sources
Authority Read relevant catalog and metrics; no resource provisioning
Budget Allocated calls/tokens and a deadline earlier than the parent deadline
Output Itemized estimate, units, source dates, assumptions and uncertainties
Stop condition Required items evaluated or explicit blocked outcome
Artifact identity Immutable result ID plus input revision

Identity and permissions must come from authenticated application state. A message saying “the supervisor authorized this” is not a credential. The worker cannot enlarge the original caller's scope by asking another agent to execute the action.

Use typed task states such as pending, running, succeeded, failed, cancelled and unknown. Distinguish “no evidence found” from “search did not finish.” A worker's final natural-language answer is not a substitute for a reliable completion record.

Coordinate shared state deliberately

The shared-blackboard pattern lets participants contribute to a common store. It does not require every agent to receive or overwrite every record.

State Recommended ownership for this example Concurrency treatment
Worker scratch context One worker Keep local and bounded
Evidence artifacts Producing worker; immutable after publication Append by unique artifact ID
Task status Scheduler/worker under a defined contract Conditional version updates
Final proposal One integration owner Publish a new version after validation
Budget counters Runtime budget service Atomic reservation and settlement

Locks can serialize access but introduce contention, expiry and recovery problems. Optimistic concurrency detects stale updates; an append-only event log records changes; a deterministic reducer can merge compatible contributions. None resolves a semantic disagreement by itself.

If two workers report different current dependency versions, keep their source revisions and timestamps. Re-read the authoritative catalog or mark the discrepancy. Last-write-wins can silently erase correct evidence, and model voting is not an authoritative database lookup.

Tip: A framework's “private state channel” can mean an internal schema, not a confidentiality boundary. Check stream and trace outputs as well as return values. LangGraph documents that full state streaming may expose private channels unless restricted. State streaming behavior.

Calculate latency and cost separately

Assume three independent investigations take 4, 6 and 8 seconds. Planning takes 2 seconds and integration takes 3 seconds. Ignore queuing and retries for this illustration.

Schedule Calculation Elapsed time
Sequential 2 + 4 + 6 + 8 + 3 23 s
Parallel 2 + max(4, 6, 8) + 3 13 s

The latency improvement is about 43%. It is not a threefold reduction, because planning, integration and the slowest worker remain on the critical path. Parallelization still performs 18 worker-seconds of work and may require more capacity at once.

If each worker consumes 2,000 input and 500 output tokens, and coordination adds 1,000 input and 300 output tokens, total model usage is 7,000 input and 1,800 output tokens. Apply the actual per-model rates separately. Duplicated context, retries, tools, storage and idle reservations add cost. Compare with a measured single-agent baseline at similar task quality.

Reserve budget before spawning children. Letting each child independently read “$1 remaining” and spend it creates an overspend race. Cancellation must reach active workers; the system should also define what happens when a remote task has already performed its side effect.

Find flaws and improve the design

Failure What to inspect Targeted improvement Tradeoff
Workers solve the wrong subtasks Assignment versus user objective Validate decomposition and dependencies before dispatch More planning work
Everyone searches the same sources Overlap in queries/artifacts Allocate distinct scope and share compact evidence IDs Less independent exploration
One worker stalls Deadline and progress record Return explicitly partial output or replace only that task Replacement can duplicate reads
Supervisor misses a contradiction Input revisions and assumptions Structured comparison and source-level verification Extra integration latency
Endless handoffs Repeated agent/task states Bound transitions and return ownership to a coordinator Some complex tasks stop early
False consensus Shared sources/models/errors Independent evidence checks and outcome tests More evaluation effort
Child repeats an external write Operation ID and status Durable idempotency and reconciliation Requires downstream support

Debate can expose mistakes, but agreement does not establish truth. A reviewer with the same model and evidence may repeat the same error. For numerical outputs, use recomputation; for code, use appropriate tests; for citations, inspect the supporting source. Record unresolved findings rather than asking agents to debate until they agree.

Choose framework and protocol boundaries

Separate the workflow shape from the tool used to implement it. LangGraph offers explicit state and transitions. Other agent frameworks provide their own handoff, workflow or hierarchy abstractions. Evaluate checkpoint behavior, cancellation, streaming exposure, version migrations and observability using the exact supported release; stars and unsupported market-share claims do not answer these questions.

For separately operated agent services, A2A can standardize discovery and task/artifact exchange. Cross-runtime communication was already possible through ordinary APIs before A2A. A protocol reduces integration differences; it does not make remote agents trustworthy or provide shared transactions. A same-team service can benefit from a protocol, and a different team can expose a conventional API. Choose according to the actual boundary. A2A specification.

Interview practice

Q1: When is one agent better than a team?

When the task has tightly coupled steps, fits one context and tool scope, and a team adds no measured quality or latency benefit. Start with the simpler baseline and add separation where it addresses an observed constraint.

Q2: Is a graph workflow necessarily multi-agent or acyclic?

Neither. Nodes may be ordinary code, tool calls or model calls. A graph may contain loops; only a DAG is acyclic. The important choices are control flow, state ownership and recovery semantics.

Q3: How would you prevent workers from overwriting each other?

Give workers separate immutable artifacts and one owner for the final result. Use conditional updates for shared status and explicit merge rules where concurrent contributions are valid. Add locks only where exclusive access is actually needed.

Q4: What is the major supervisor risk?

Bad decomposition can produce convincing answers to the wrong questions. Check task coverage and dependencies before dispatch, and validate worker evidence against the original goal during integration. Preserve enough source context to challenge a worker's conclusion.

Q5: Does parallelism make the task cheaper?

It can reduce elapsed time, but total work and model usage may rise. Calculate critical-path latency and sum each worker's resource usage separately. Include coordination, duplicate context and retries.

Q6: How do you evaluate the team?

Measure final task success, evidence quality, permission violations, cost and tail latency. Diagnose decomposition, individual workers and integration separately. Compare to a single-agent baseline and remove workers that add no benefit.

Q7: How do you close this design in an interview?

State the chosen pattern, why the subtasks are independent, who owns shared state and how failures terminate. Explain the budget and critical path, then identify the next experiment: whether parallel investigations improve success or latency enough to cover integration and operating cost.

Final notes

Recall card: Divide by responsibility → delegate with scope → return evidence → reconcile → verify. More participants create more coordination obligations; agent count is not a measure of design quality.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Tool Use and MCP
NEXT LESSONAgent Memory and State →

Explore the diagram