Framework selection is the choice of reusable software that fits an application's execution, data, and operational requirements. It is not a ranking of which library is most advanced. The right choice depends on the behavior you need, the capabilities you can verify, and the work your team can maintain.
A useful interview answer explains a decision that another engineer could reproduce. “Use DSPy for 99% reliability” or “use two frameworks because the system is complex” does not provide that evidence. Reliability is an observed property of the whole system under specified conditions.
Separate the categories
| Category | Main job | Examples to evaluate | Selection question |
|---|---|---|---|
| Model client or gateway | Call models and normalize selected interfaces | Provider SDKs, an application-owned adapter, a gateway | Which differences should be abstracted and which must stay visible? |
| Agent runtime/orchestration | Run tools, state transitions, handoffs, and loops | LangGraph, Microsoft Agent Framework, CrewAI, agent SDKs | How does execution stop, resume, and recover? |
| Retrieval/data framework | Ingest, index, and retrieve evidence | LlamaIndex or targeted retrieval components | Does it improve this corpus's quality and lifecycle management? |
| Program optimization | Search instructions, examples, or supported trainable components | DSPy | Is the evaluation objective good enough to optimize? |
| Observability/evaluation | Inspect execution and measure behavior | LangSmith or another suitable tracing/evaluation stack | Can we diagnose failures and compare releases? |
| Coding product or coding-agent SDK | Operate on repositories and development tools | Claude Code, Cline, Cursor, OpenHands, other coding products | Where does code execute and how are changes reviewed? |
| Hosted service/UI integration | Operate a runtime or provide a product interface | A managed agent service, ChatKit, plugin integration | What does the service own, and what remains ours? |
These categories overlap. A framework can provide both retrieval and workflows. Avoid invented “L1/L2/L3” tiers unless you explicitly define them for a specific diagram; they are not standard maturity levels.
Write the requirements first
For a learner-facing interview coach, an initial selection brief could be:
Functional requirements
- Retrieve relevant teaching material with source links.
- Review a submitted answer against a versioned rubric.
- Preserve progress across sessions.
- Let the learner inspect and correct generated feedback.
- Optionally hand off scheduling requests to an existing booking service.
Non-functional requirements
- Enforce learner and course-access boundaries.
- Meet separate latency targets for interactive help and longer reviews.
- Bound cost per completed review and handle provider limits.
- Recover persisted work without duplicating external actions.
- Export useful diagnostics and support a tested rollback.
- Fit the team's actual language, deployment environment, and maintenance capacity.
Agree on measurable targets with the interviewer. State assumptions explicitly rather than assigning universal thresholds. For example, a proposed ten-second feedback target is a product assumption to validate, not an inherent requirement of all interview coaches.
Begin with the smallest sufficient design
Read diagram source
flowchart TD
A[Define task and constraints] --> B{Can ordinary code solve it?}
B -->|Yes| C[Use ordinary code]
B -->|No| D[One model call or bounded tool loop]
D --> E[Evaluate representative failures]
E --> F{Which capability is missing?}
F -->|Data quality or retrieval| G[Evaluate retrieval components]
F -->|State and recovery| H[Evaluate workflow runtime]
F -->|Measured prompt quality| I[Evaluate optimization]
F -->|No material gap| J[Keep the baseline]
G --> K[Compare total behavior and operating cost]
H --> K
I --> K
K --> L[Select and record an exit plan]
A thin implementation still needs timeouts, validation, telemetry, and error handling. A framework may reduce repeated work in those areas. Conversely, a framework that exposes ten features does not require using all ten.
Use a capability matrix, then test the claims
The following rows identify starting points for investigation. They are not exclusive assignments or performance rankings. The linked lessons include current primary documentation.
| Candidate | Capability worth evaluating | Proof to request before choosing |
|---|---|---|
| LangGraph | Explicit state, transitions, checkpoints, and interrupts | Crash/resume, state migration, and effect-reconciliation exercise |
| LlamaIndex | Document ingestion, indexing, query engines, and workflows | Corpus-specific retrieval tests and revision/deletion handling |
| DSPy | Composable model programs and optimization | Held-out improvement under a fixed search and runtime budget |
| CrewAI | Task-oriented crews inside controlled flows | Routing, partial failure, termination, and persistence tests |
| Microsoft Agent Framework | Agent/session/workflow integration | Target-language support and compatibility with existing services |
| Claude Agent SDK | Embedded coding/tool loop | Deployment, permission, sandbox, and recovery behavior |
| OpenAI Agents SDK | Application-run agents, tools, and handoffs | Guardrail boundaries, continuation, and storage integration |
| Google ADK | Agents and workflow/deployment integrations | The exact feature in the selected language and hosting mode |
Read lifecycle notices as part of selection. For example, AutoGen is in maintenance mode, and Agent Builder has an announced shutdown. These have different implications; neither is well represented by a generic “enterprise ready” score. See the current SDK landscape.
Compare three realistic candidates
For the coach, shortlist:
- Application code plus model client: a fixed retrieval/review sequence with ordinary service persistence.
- A retrieval framework plus application workflow: useful if document processing and query composition dominate development work.
- A stateful agent/workflow framework: useful if tasks branch, suspend, resume, or require sophisticated recovery.
Run the same tasks, model configuration where possible, evidence, and release criteria for all candidates. Include ordinary requests, missing evidence, malformed outputs, provider timeouts, cancellation, and access changes.
| Dimension | Measurement | Common misleading substitute |
|---|---|---|
| Quality | Task success and error severity on held-out cases | A polished demo response |
| Latency | End-to-end p50/p95 with concurrency and timeout rates | One warm model call |
| Cost | Total spend per completed task, including failed attempts | Price per token alone |
| Recovery | Correct outcomes after injected failures | A “supports persistence” checkbox |
| Security | Tests at actual data/action boundaries | Presence of a guardrail API |
| Maintainability | Time to diagnose, upgrade, and change one real flow | GitHub stars or marketing adoption claims |
| Portability | A working alternate provider/runtime path | A shared method name |
Interview tip: distinguish hard requirements from preferences. A candidate that cannot meet a required data boundary should not win because a weighted average gives it a high convenience score.
After eliminating infeasible choices, a weighted comparison can make preferences explicit. The weights and scores are judgments, not objective properties of libraries. Record the evidence behind them and check whether a small weight change reverses the decision.
Count the full cost
Use an explicit model:
total cost = implementation + maintenance + infrastructure + model/tool usage + observability + migration
Treat uncertain incident and vendor-change exposure separately rather than assigning fabricated precise dollar values.
An illustrative first-year comparison:
| Cost assumption | Thin custom path | Framework-assisted path |
|---|---|---|
| Initial engineering at an assumed $100/hour | 100 hours = $10,000 | 40 hours = $4,000 |
| Monthly maintenance at that rate | 10 hours = $1,000 | 4 hours = $400 |
| Additional monthly platform expense | $0 | $300 |
| First-year subtotal for these components | $22,000 | $12,400 |
This example favors the framework only under the stated assumptions. It excludes common hosting and model costs. If an abstraction causes expensive retries or makes debugging much harder, the result can reverse. Obtain real estimates from the representative implementation rather than treating this table as vendor pricing.
Also compare costs per successful task. If two systems each spend $100, but one completes 80 tasks and the other 100, those observed costs are $1.25 and $1.00 per success. State whether success means an accepted answer, a confirmed business operation, or something else.
Build, adopt a framework, or buy a service?
| Choice | Benefit | Responsibility retained | Exit question |
|---|---|---|---|
| Build a thin layer | Direct control and few abstractions | Implement and maintain every required capability | Is the custom surface small enough to sustain? |
| Adopt a library/framework | Reuse components and execution machinery | Deployment, domain correctness, versioning, operations | Can application contracts survive a library replacement? |
| Buy a managed service | Reduce selected operational work | Data policy, integration correctness, product behavior | Can state, traces, and artifacts be exported usefully? |
| Combine approaches | Keep critical boundaries owned while outsourcing others | Integration and failure handling between components | Are there too many runtimes for the value delivered? |
“Managed” does not mean all storage and tools are hosted. “Open source” does not mean every platform feature is free or that any model works well. Inspect the specific component and deployment option.
For example, OpenHands documentation distinguishes its software-agent components, browser client, managed service, and enterprise offerings. A blanket statement that it always requires self-hosting is inaccurate. Likewise, a coding product should be assessed as a particular interface and execution environment, not just a brand name.
Keep a practical exit path
For a multi-provider design, own the contracts that matter:
- Domain operations with typed requests, authorization, stable operation IDs, and explicit results.
- Evaluation datasets, scoring rules, and acceptance thresholds.
- Source content and its revision/permission metadata.
- Exportable conversation/workflow state where feasible.
- Traces and cost records tied to application task IDs.
- A tested fallback and rollback procedure.
A model adapter can normalize request shapes. It cannot make different models equally capable or reproduce provider-specific tools. A fallback may also move data to a different service; enforce the same application requirements before allowing it.
Use MCP or A2A when they solve a real integration boundary. For a single in-process service, an ordinary function may be simpler. Protocol compatibility does not remove differences in authentication, state, capabilities, or runtime behavior.
Keep coding tools distinct from the product runtime
Cline, Cursor, and other coding products can help build or operate on a repository. Some also expose SDKs, command-line interfaces, or remote execution. Decide which surface you are evaluating. The product once documented at Windsurf's entry point now has Devin Desktop documentation; verify current names and lifecycle instead of preserving an old brand comparison indefinitely.
For coding automation, evaluate repository isolation, command/network permissions, secret access, meaningful tests, change review, and reproducible environments. “Works with any LLM” is not a realistic promise of equal capability. The tool chosen to build an application need not be the runtime used to serve that application.
Decision record and interview practice
Close the selection with a compact record:
| Field | What to write |
|---|---|
| Decision | Selected approach and the problem it solves |
| Alternatives | Real candidates tested, including the baseline |
| Evidence | Quality, recovery, latency, cost, and maintenance results |
| Accepted costs | Specific limitations the team can tolerate |
| Review trigger | A measurable change that warrants reconsideration |
| Exit path | Contracts, data, and state needed to migrate |
- Why not use the most feature-rich framework? Unused features do not solve a requirement and can add dependencies or complexity.
- Does DSPy guarantee 99% reliability? No. Optimization can improve a chosen empirical metric; the whole system must be evaluated under defined conditions.
- How do you choose between a retrieval framework and a workflow framework? Identify the missing capability, test representative failures, and account for overlap rather than applying a rigid category rule.
- What evidence proves portability? A tested alternate configuration meeting the required behavior, not just a common API.
- When does managed hosting make sense? When its operational benefits exceed its cost and constraints, with acceptable data handling and an exit plan.
- Should every system implement MCP and A2A? No. Adopt a protocol for a concrete interoperability need.
- What is the strongest reason to change a working framework? A demonstrated requirement, support issue, or operating cost that the current choice cannot reasonably address.
Final notes
Remember requirements → baseline → failure tests → total cost → decision record. A defensible selection names what the framework solves, what remains application work, and what evidence would change the decision.
Continue with coding-agent workflows or managing framework change.