Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Multi-agent frameworks: CrewAI, AutoGen, and current SDKs

By Anup Rai8 min readReviewed September 2026

A multi-agent framework coordinates multiple model-driven components, their tools, and their shared work. It can help organize delegation, state, and execution. It cannot make several agents' answers independent, truthful, or useful merely by giving them different role names.

In a Learnastra design discussion, start with one agent or an ordinary workflow. Add a specialist only when it contributes a distinct capability, context boundary, permission boundary, or measurable quality improvement. The multi-agent orchestration lesson develops the underlying patterns; this chapter connects those patterns to actual frameworks.

Separate the orchestration patterns

Pattern Who owns the next step? Example Main tradeoff
Sequential workflow Application-defined order Extract requirements, then produce feedback Predictable but less flexible
Manager with specialist tools Manager retains the final response Ask a capacity reviewer for a bounded calculation Extra coordination and a manager bottleneck
Handoff Control passes to a specialist Route a billing conversation to a billing agent History, authority, and ownership must transfer correctly
Parallel specialists Coordinator joins independent results Review latency and storage assumptions separately More work and a policy for partial failures
Peer exchange or swarm Participants route work under allowed rules Specialists exchange findings Harder termination, routing, and debugging

A swarm is an orchestration pattern, not a synonym for every multi-agent application. Parallel calls do not imply peer-to-peer communication. Handoffs and agents-as-tools have different ownership semantics; see the OpenAI orchestration guide for a concrete SDK example.

CrewAI: agents, tasks, crews, and flows

CrewAI's agent abstraction describes a worker's role, goal, tools, and behavior. A task describes work and its expected output. A crew groups agents and tasks with a process. A Flow provides event-driven control and state around work that can include crews or ordinary application functions.

The documented crew process types are sequential and hierarchical. In a hierarchical process, a configured manager model or manager agent coordinates delegation. Do not assume that a “consensual” process is a supported enum because an older comparison lists it. See the current process documentation.

Flows use entry points, listeners, and conditional routing. State can be structured, and persistence is configurable. These are useful mechanisms for an application-controlled outer workflow; they do not establish atomicity with an external database or API. See Flows.

A grounded example: review a practice design

Functional requirements:

  1. Accept a learner's design and the selected interview rubric.
  2. Identify stated requirements before evaluating choices.
  3. Check numerical capacity assumptions with ordinary code.
  4. Produce evidence-based feedback with links to relevant concepts.
  5. Return partial feedback explicitly when a specialist fails.

Non-functional requirements:

  1. Enforce one shared deadline and cost budget for the entire review.
  2. Keep answers isolated by learner and session.
  3. Track every finding to a source passage or calculation.
  4. Avoid treating agreement among agents as proof.
  5. Preserve the original answer and the versions used for the review.

A simple baseline runs extraction, deterministic checks, and one feedback model in sequence. If evaluations show that separating capacity and reliability review improves feedback, introduce those specialists inside an explicit outer flow.

Architecture / visual model
flowchart TD A[Answer and rubric version] --> B[Extract stated requirements] B --> C[Run deterministic calculations] C --> D[Bounded specialist review] D --> E[Capacity findings with evidence] D --> F[Reliability findings with evidence] E --> G[Join by answer revision] F --> G G --> H{Required results available?} H -->|No| I[Label partial result or retry within budget] H -->|Yes| J[Validate and synthesize feedback] J --> K[Learner reviews feedback]
Read diagram source
flowchart TD
    A[Answer and rubric version] --> B[Extract stated requirements]
    B --> C[Run deterministic calculations]
    C --> D[Bounded specialist review]
    D --> E[Capacity findings with evidence]
    D --> F[Reliability findings with evidence]
    E --> G[Join by answer revision]
    F --> G
    G --> H{Required results available?}
    H -->|No| I[Label partial result or retry within budget]
    H -->|Yes| J[Validate and synthesize feedback]
    J --> K[Learner reviews feedback]

The specialists need not be two model instances. One may be a pure calculation service. Assigning the title “senior architect” to an agent does not improve its mathematical reliability.

Repair the baseline without adding uncontrolled conversation

Observed problem Change Benefit Cost or limitation
Capacity numbers are repeatedly wrong Deterministic arithmetic with explicit units Reproducible calculations Inputs and assumptions still need review
Specialists grade different answer versions Key results by answer and rubric revision Prevents invalid joins Revision tracking
A reviewer fails but the response appears complete Typed completion/partial/failure states Honest, usable feedback More UI states
Reviewers repeat the same criticism Deduplicate by claim and evidence Less distracting feedback Deduplication can merge distinct issues
The manager keeps asking for more reviews Shared hard budget and measurable stop condition Bounded spend and latency Some requests stop incomplete
Resumed work repeats a publication step Idempotent publication and operation reconciliation Prevents duplicate side effects Additional durable records

Crew state persistence does not prove that a publication or other external effect completed exactly once. Test recovery at the boundary between a successful external operation and the next saved state. Also distinguish open-source runtime features from hosted-platform controls such as SSO or workspace administration; buying a platform feature does not automatically apply that policy to every custom tool.

AutoGen and Microsoft Agent Framework

As checked in September 2026, the official AutoGen repository describes maintenance mode, with no new features and community management. Existing code is not automatically unusable. The project recommends Microsoft Agent Framework for new development and provides a migration path. Do not describe maintenance mode as an immediate shutdown or promise a particular security-patch service level. See the AutoGen repository notice.

The important migration questions are behavioral:

  1. How do agents receive model clients and tools?
  2. How are messages, sessions, and streaming events represented?
  3. Which component decides who speaks or acts next?
  4. Where are termination, authorization, and persisted recovery enforced?
  5. What happens to existing sessions during rollout?

Use Microsoft's AutoGen migration guide for the installed target version. An old GroupChat configuration does not become an equivalent production workflow through a class-name substitution. See Semantic Kernel and Agent Framework for a domain-service design and migration checklist.

The current SDK landscape

This is a capability map, not a ranking. Confirm the specific language package, provider integration, and support status required for deployment.

Framework or SDK Useful starting surface What the application must still decide
CrewAI Agents/tasks/crews and an outer Flow Process, validation, runtime limits, and effect recovery
Microsoft Agent Framework Agents, sessions, middleware, and workflows Domain policy, persistence configuration, and migration compatibility
LangGraph State transitions, reducers, checkpoints, and interrupts Graph semantics, authorization, and external-effect safety
Claude Agent SDK Claude Code's tool loop embedded in Python or TypeScript Host process, permissions, isolation, and tool scope
OpenAI Agents SDK Agents, tools, handoffs, guardrails, and tracing Deployment, storage, review policy, and service integration
Google ADK Agents, workflows, tools, evaluation, and deployment integrations Language-specific feature support, runtime choice, and domain controls

The Claude Agent SDK embeds the Claude Code agent machinery in a process you operate. It is distinct from the interactive CLI and from a basic model client SDK. Built-in file and command tools increase what the application can do; configure their permissions and execution environment accordingly.

The OpenAI Agents SDK runs the agent loop in your application. Guardrails have specific boundaries: agent-level input/output checks are not automatically checks around every internal tool call. Put critical validation at the actual effect boundary and review the guardrail documentation.

Google ADK offers multiple language implementations and integrations. Do not infer identical feature coverage from the list of supported languages. A managed deployment option is different from a requirement to host every ADK application on one cloud.

Keep product surfaces separate

Older material may group several OpenAI products under AgentKit. Identify the component actually used:

Surface Purpose Current design implication
Agents SDK Application-side orchestration Your service owns the surrounding runtime
ChatKit An embeddable chat interface Connect it to a supported server-side implementation
Agent Builder Visual workflow authoring Deprecated; shutdown scheduled for November 30, 2026
Plugin/MCP UI integration Expose tools and optional interactive UI inside supported clients Different from an application's agent runtime

The Agent Builder deprecation notice dates its announcement to June 3, 2026. ChatKit remains available; new integrations should follow its current server integration guidance. The former Apps SDK documentation entry now routes to plugin documentation. Do not mistake a plugin UI, an agent runner, and a hosted workflow service for interchangeable products.

Interoperability and termination

MCP and A2A address different integration boundaries. MCP connects clients to server capabilities such as tools and resources. A2A supports interaction with remote agents and task-oriented work. Neither is required when an ordinary local function or existing service API is sufficient. Neither standardizes all model behavior or application policy.

A handoff must carry the minimum necessary task context, allowed authority, completion contract, and budget. Use explicit task IDs and cancellation/deadline handling across remote boundaries. Treat peer messages as untrusted inputs; a remote agent's claim that it is authorized is not proof.

For loop control, use concrete signals:

  1. Maximum elapsed time and total spend.
  2. Maximum model/tool calls and bounded retries.
  3. A task-specific completion condition.
  4. Repeated identical actions or repeated failures without new evidence.
  5. A defined incomplete-result or escalation path.

A critic model can contribute a signal, but it must not be the only mechanism capable of stopping the system. “100,000 tokens in two minutes” is not a universal safe threshold.

Latency/cost example: two independent reviewers take 3 and 5 seconds, and synthesis takes 2 seconds. Ignoring overhead, parallel review gives max(3,5)+2 = 7 seconds instead of 10 seconds sequentially. It still executes both reviews and synthesis. Parallelism can reduce elapsed time while leaving total model work unchanged or increasing it through coordination.

Interview questions and answer notes

  1. Does giving agents different roles create independent evidence? No. They may use the same model, context, and mistaken assumptions.
  2. When is a specialist a tool rather than a handoff target? When the outer agent retains responsibility for the reply and needs a bounded result.
  3. Does CrewAI currently provide a consensual process enum? The referenced documentation lists sequential and hierarchical processes. Verify supported APIs instead of relying on an old comparison.
  4. What does AutoGen maintenance mode mean for a live system? Review support exposure and plan a tested migration; do not assume immediate failure or guaranteed ongoing support.
  5. Why can a parallel review be faster but not cheaper? Both branches still consume resources; only their waiting periods overlap.
  6. Does adding MCP eliminate vendor lock-in? It can standardize a tool boundary, but model behavior, state, deployment, and application semantics remain dependencies.
  7. What stops an agent team that keeps debating? Enforced global budgets and an explicit termination path, supplemented by progress checks.
  8. Would you start a new production dependency on Agent Builder? Its announced shutdown makes it an unsuitable long-term foundation; evaluate the supported runtime and UI alternatives separately.

Final notes

Remember roles organize work; contracts make it checkable; budgets keep it bounded. Choose one clear owner for the response, test partial failures, and require evidence before expanding a single-agent baseline into a team.

Next: Choosing an AI framework.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Semantic Kernel and Microsoft Agent Framework
NEXT LESSONChoosing an AI framework: requirements, evidence, and operating cost →

Explore the diagram