A multi-agent framework coordinates multiple model-driven components, their tools, and their shared work. It can help organize delegation, state, and execution. It cannot make several agents' answers independent, truthful, or useful merely by giving them different role names.
In a Learnastra design discussion, start with one agent or an ordinary workflow. Add a specialist only when it contributes a distinct capability, context boundary, permission boundary, or measurable quality improvement. The multi-agent orchestration lesson develops the underlying patterns; this chapter connects those patterns to actual frameworks.
Separate the orchestration patterns
| Pattern | Who owns the next step? | Example | Main tradeoff |
|---|---|---|---|
| Sequential workflow | Application-defined order | Extract requirements, then produce feedback | Predictable but less flexible |
| Manager with specialist tools | Manager retains the final response | Ask a capacity reviewer for a bounded calculation | Extra coordination and a manager bottleneck |
| Handoff | Control passes to a specialist | Route a billing conversation to a billing agent | History, authority, and ownership must transfer correctly |
| Parallel specialists | Coordinator joins independent results | Review latency and storage assumptions separately | More work and a policy for partial failures |
| Peer exchange or swarm | Participants route work under allowed rules | Specialists exchange findings | Harder termination, routing, and debugging |
A swarm is an orchestration pattern, not a synonym for every multi-agent application. Parallel calls do not imply peer-to-peer communication. Handoffs and agents-as-tools have different ownership semantics; see the OpenAI orchestration guide for a concrete SDK example.
CrewAI: agents, tasks, crews, and flows
CrewAI's agent abstraction describes a worker's role, goal, tools, and behavior. A task describes work and its expected output. A crew groups agents and tasks with a process. A Flow provides event-driven control and state around work that can include crews or ordinary application functions.
The documented crew process types are sequential and hierarchical. In a hierarchical process, a configured manager model or manager agent coordinates delegation. Do not assume that a “consensual” process is a supported enum because an older comparison lists it. See the current process documentation.
Flows use entry points, listeners, and conditional routing. State can be structured, and persistence is configurable. These are useful mechanisms for an application-controlled outer workflow; they do not establish atomicity with an external database or API. See Flows.
A grounded example: review a practice design
Functional requirements:
- Accept a learner's design and the selected interview rubric.
- Identify stated requirements before evaluating choices.
- Check numerical capacity assumptions with ordinary code.
- Produce evidence-based feedback with links to relevant concepts.
- Return partial feedback explicitly when a specialist fails.
Non-functional requirements:
- Enforce one shared deadline and cost budget for the entire review.
- Keep answers isolated by learner and session.
- Track every finding to a source passage or calculation.
- Avoid treating agreement among agents as proof.
- Preserve the original answer and the versions used for the review.
A simple baseline runs extraction, deterministic checks, and one feedback model in sequence. If evaluations show that separating capacity and reliability review improves feedback, introduce those specialists inside an explicit outer flow.
Read diagram source
flowchart TD
A[Answer and rubric version] --> B[Extract stated requirements]
B --> C[Run deterministic calculations]
C --> D[Bounded specialist review]
D --> E[Capacity findings with evidence]
D --> F[Reliability findings with evidence]
E --> G[Join by answer revision]
F --> G
G --> H{Required results available?}
H -->|No| I[Label partial result or retry within budget]
H -->|Yes| J[Validate and synthesize feedback]
J --> K[Learner reviews feedback]
The specialists need not be two model instances. One may be a pure calculation service. Assigning the title “senior architect” to an agent does not improve its mathematical reliability.
Repair the baseline without adding uncontrolled conversation
| Observed problem | Change | Benefit | Cost or limitation |
|---|---|---|---|
| Capacity numbers are repeatedly wrong | Deterministic arithmetic with explicit units | Reproducible calculations | Inputs and assumptions still need review |
| Specialists grade different answer versions | Key results by answer and rubric revision | Prevents invalid joins | Revision tracking |
| A reviewer fails but the response appears complete | Typed completion/partial/failure states | Honest, usable feedback | More UI states |
| Reviewers repeat the same criticism | Deduplicate by claim and evidence | Less distracting feedback | Deduplication can merge distinct issues |
| The manager keeps asking for more reviews | Shared hard budget and measurable stop condition | Bounded spend and latency | Some requests stop incomplete |
| Resumed work repeats a publication step | Idempotent publication and operation reconciliation | Prevents duplicate side effects | Additional durable records |
Crew state persistence does not prove that a publication or other external effect completed exactly once. Test recovery at the boundary between a successful external operation and the next saved state. Also distinguish open-source runtime features from hosted-platform controls such as SSO or workspace administration; buying a platform feature does not automatically apply that policy to every custom tool.
AutoGen and Microsoft Agent Framework
As checked in September 2026, the official AutoGen repository describes maintenance mode, with no new features and community management. Existing code is not automatically unusable. The project recommends Microsoft Agent Framework for new development and provides a migration path. Do not describe maintenance mode as an immediate shutdown or promise a particular security-patch service level. See the AutoGen repository notice.
The important migration questions are behavioral:
- How do agents receive model clients and tools?
- How are messages, sessions, and streaming events represented?
- Which component decides who speaks or acts next?
- Where are termination, authorization, and persisted recovery enforced?
- What happens to existing sessions during rollout?
Use Microsoft's AutoGen migration guide for the installed target version. An old GroupChat configuration does not become an equivalent production workflow through a class-name substitution. See Semantic Kernel and Agent Framework for a domain-service design and migration checklist.
The current SDK landscape
This is a capability map, not a ranking. Confirm the specific language package, provider integration, and support status required for deployment.
| Framework or SDK | Useful starting surface | What the application must still decide |
|---|---|---|
| CrewAI | Agents/tasks/crews and an outer Flow | Process, validation, runtime limits, and effect recovery |
| Microsoft Agent Framework | Agents, sessions, middleware, and workflows | Domain policy, persistence configuration, and migration compatibility |
| LangGraph | State transitions, reducers, checkpoints, and interrupts | Graph semantics, authorization, and external-effect safety |
| Claude Agent SDK | Claude Code's tool loop embedded in Python or TypeScript | Host process, permissions, isolation, and tool scope |
| OpenAI Agents SDK | Agents, tools, handoffs, guardrails, and tracing | Deployment, storage, review policy, and service integration |
| Google ADK | Agents, workflows, tools, evaluation, and deployment integrations | Language-specific feature support, runtime choice, and domain controls |
The Claude Agent SDK embeds the Claude Code agent machinery in a process you operate. It is distinct from the interactive CLI and from a basic model client SDK. Built-in file and command tools increase what the application can do; configure their permissions and execution environment accordingly.
The OpenAI Agents SDK runs the agent loop in your application. Guardrails have specific boundaries: agent-level input/output checks are not automatically checks around every internal tool call. Put critical validation at the actual effect boundary and review the guardrail documentation.
Google ADK offers multiple language implementations and integrations. Do not infer identical feature coverage from the list of supported languages. A managed deployment option is different from a requirement to host every ADK application on one cloud.
Keep product surfaces separate
Older material may group several OpenAI products under AgentKit. Identify the component actually used:
| Surface | Purpose | Current design implication |
|---|---|---|
| Agents SDK | Application-side orchestration | Your service owns the surrounding runtime |
| ChatKit | An embeddable chat interface | Connect it to a supported server-side implementation |
| Agent Builder | Visual workflow authoring | Deprecated; shutdown scheduled for November 30, 2026 |
| Plugin/MCP UI integration | Expose tools and optional interactive UI inside supported clients | Different from an application's agent runtime |
The Agent Builder deprecation notice dates its announcement to June 3, 2026. ChatKit remains available; new integrations should follow its current server integration guidance. The former Apps SDK documentation entry now routes to plugin documentation. Do not mistake a plugin UI, an agent runner, and a hosted workflow service for interchangeable products.
Interoperability and termination
MCP and A2A address different integration boundaries. MCP connects clients to server capabilities such as tools and resources. A2A supports interaction with remote agents and task-oriented work. Neither is required when an ordinary local function or existing service API is sufficient. Neither standardizes all model behavior or application policy.
A handoff must carry the minimum necessary task context, allowed authority, completion contract, and budget. Use explicit task IDs and cancellation/deadline handling across remote boundaries. Treat peer messages as untrusted inputs; a remote agent's claim that it is authorized is not proof.
For loop control, use concrete signals:
- Maximum elapsed time and total spend.
- Maximum model/tool calls and bounded retries.
- A task-specific completion condition.
- Repeated identical actions or repeated failures without new evidence.
- A defined incomplete-result or escalation path.
A critic model can contribute a signal, but it must not be the only mechanism capable of stopping the system. “100,000 tokens in two minutes” is not a universal safe threshold.
Latency/cost example: two independent reviewers take 3 and 5 seconds, and synthesis takes 2 seconds. Ignoring overhead, parallel review gives max(3,5)+2 = 7 seconds instead of 10 seconds sequentially. It still executes both reviews and synthesis. Parallelism can reduce elapsed time while leaving total model work unchanged or increasing it through coordination.
Interview questions and answer notes
- Does giving agents different roles create independent evidence? No. They may use the same model, context, and mistaken assumptions.
- When is a specialist a tool rather than a handoff target? When the outer agent retains responsibility for the reply and needs a bounded result.
- Does CrewAI currently provide a consensual process enum? The referenced documentation lists sequential and hierarchical processes. Verify supported APIs instead of relying on an old comparison.
- What does AutoGen maintenance mode mean for a live system? Review support exposure and plan a tested migration; do not assume immediate failure or guaranteed ongoing support.
- Why can a parallel review be faster but not cheaper? Both branches still consume resources; only their waiting periods overlap.
- Does adding MCP eliminate vendor lock-in? It can standardize a tool boundary, but model behavior, state, deployment, and application semantics remain dependencies.
- What stops an agent team that keeps debating? Enforced global budgets and an explicit termination path, supplemented by progress checks.
- Would you start a new production dependency on Agent Builder? Its announced shutdown makes it an unsuitable long-term foundation; evaluate the supported runtime and UI alternatives separately.
Final notes
Remember roles organize work; contracts make it checkable; budgets keep it bounded. Choose one clear owner for the response, test partial failures, and require evidence before expanding a single-agent baseline into a team.
Next: Choosing an AI framework.