Pydantic AI is a Python framework for building model-driven applications with typed dependencies, tools, and outputs. Mastra is a TypeScript framework with agents, tools, workflows, storage integrations, and development tooling. Both can reduce integration work. Neither turns a model's prediction into a verified fact merely by parsing it.
The interview question is: which contracts should the application enforce, and which framework makes those contracts easier to implement and test? Language fit matters, but there is no universal rule that one framework is for simple agents and another is for all complex workflows.
What “typed” actually guarantees
| Boundary | What it can check | What it cannot establish by itself |
|---|---|---|
| Static type checking | Compatible function arguments and result types during development | The contents of a network response at runtime |
| Runtime schema validation | Required fields, supported types, configured ranges and enums | Whether a plausible value is factually correct |
| Domain validation | Topic exists; durations fit the study budget | Whether the plan is educationally effective |
| Authorization | Authenticated learner can access this module | Whether the model's explanation is useful |
| Behavioral evaluation | Quality on representative cases | Perfect behavior on every future input |
For example, { "topicId": "made-up-topic", "minutes": 30 } can satisfy a schema containing a string and a positive integer. A catalog lookup must reject the unknown topic. A real topic may still be unsuitable for a learner who has not studied its prerequisites.
Validation behavior also depends on configuration: some validators coerce values or ignore unknown fields. Do not assume “typed” always means strict rejection of every extra field. Treat structured output as one layer of a complete contract.
Interview exercise: generate a weekly study plan
Functional requirements
- Accept the learner's goal, available minutes and selected modules.
- Retrieve accessible topics, prerequisites and relevant quiz results.
- Propose an ordered set of study sessions with topic IDs and durations.
- Validate the proposal and explain why each session was selected.
- Save the accepted plan and allow subsequent revision.
Non-functional requirements
- Scope quiz history and saved plans to the authenticated account.
- Reject invalid topic IDs, inaccessible modules and over-budget plans.
- Bound retries, model expenditure and response time.
- Preserve plan versions so concurrent edits do not silently overwrite work.
- Record enough evidence to diagnose failures without logging unnecessary learner data.
The model proposes a plan. Application code owns identity, access, catalog correctness, arithmetic, and persistence.
Read diagram source
flowchart TD
A[Authenticated request and time budget] --> B[Read authorized catalog and quiz summary]
B --> C[Model proposes typed study plan]
C --> D[Schema validation]
D --> E[Catalog, access and budget checks]
E -->|Valid| F[Show proposal and explanation]
E -->|Repairable, budget remains| C
E -->|Invalid or limit reached| G[Clear failure or deterministic fallback]
F --> H[Save with expected plan version]
H --> I[Versioned plan store]
Begin with a read-only proposal and an ordinary save endpoint. Add a persistent workflow only when the product requires a long wait, background work, or recovery across process failure.
Pydantic AI: output types and trusted dependencies
An agent declares its output type. Current output documentation distinguishes tool-based, provider-native, and prompted structured-output modes. Model/provider support differs; changing the model string does not prove that an existing output mode remains compatible.
This example defines a strict plan shape. It assumes installed compatible versions of Pydantic and Pydantic AI; the caller supplies a supported configured model. It is a contract example, not a complete authenticated service.
from dataclasses import dataclass
from pydantic import BaseModel, ConfigDict, Field
from pydantic_ai import Agent, ModelRetry, RunContext
class Session(BaseModel):
model_config = ConfigDict(extra="forbid", strict=True)
topic_id: str
minutes: int = Field(ge=5, le=120)
class StudyPlan(BaseModel):
model_config = ConfigDict(extra="forbid", strict=True)
sessions: list[Session] = Field(min_length=1, max_length=20)
@dataclass(frozen=True)
class PlanContext:
allowed_topic_ids: frozenset[str]
available_minutes: int
def build_planner(model):
planner = Agent(
model,
deps_type=PlanContext,
output_type=StudyPlan,
instructions="Propose a study plan using the supplied catalog and budget.",
retries=1,
)
@planner.output_validator
def validate_plan(ctx: RunContext[PlanContext], plan: StudyPlan) -> StudyPlan:
if any(s.topic_id not in ctx.deps.allowed_topic_ids for s in plan.sessions):
raise ModelRetry("Use only topics in the supplied catalog.")
if sum(s.minutes for s in plan.sessions) > ctx.deps.available_minutes:
raise ModelRetry("The plan exceeds the available study time.")
return plan
return planner
The service derives PlanContext from trusted account/catalog data, supplies an appropriate catalog in the model input, and handles exhausted validation retries. The output validator does not enforce prerequisites yet; add an explicit prerequisite rule if that is part of the product contract. Recheck access and the expected plan version when saving because permissions or state can change after generation.
deps_type and RunContext connect typed application dependencies to tools and validators. The dependency declaration aids static checking; it is not automatic runtime authentication. Do not let the model choose the account ID or construct a privileged database client.
Tools can be ordinary typed functions with schemas derived from their signatures. That reduces duplicate schema definitions, but annotations, docstrings and implementation can still disagree semantically. Test the underlying service function separately from the model's decision to call it.
Mastra: tools, structured results and workflows
Mastra's current tool API uses execute(inputData, context): validated tool arguments are the first parameter; execution metadata is the second. Older examples using a single { context } wrapper do not describe this current signature. Schemas can use supported libraries such as Zod.
A small deterministic tool illustrates the boundary without requiring model access:
import { createTool } from '@mastra/core/tools';
import { z } from 'zod';
export const totalStudyMinutes = createTool({
id: 'total-study-minutes',
description: 'Add the durations of proposed study sessions.',
inputSchema: z.object({
durations: z.array(z.number().int().min(5).max(120)).min(1).max(20),
}).strict(),
outputSchema: z.object({ totalMinutes: z.number().int() }),
execute: async ({ durations }) => ({
totalMinutes: durations.reduce((sum, minutes) => sum + minutes, 0),
}),
});
For this study planner, the application should calculate the total itself even if the agent has such a tool. Tool availability does not force the model to call it, and the model's final text can misreport its result.
Mastra supports structured agent output through a schema supplied to generation. Configure failure behavior explicitly. A fallback object must be distinguishable from a valid personalized plan; silently returning an empty “successful” plan can hide a provider or parsing outage.
Use workflows when the sequence requires explicit branching or waits. Suspend/resume stores snapshots in the configured storage provider. Recovering a suspended run does not automatically authorize whoever possesses its run ID. Authenticate the resume request and verify its account, expected state and permissible transition.
Studio, launched with the development server, helps inspect and exercise agents and workflows. An accessible development UI is not automatically an appropriate public production endpoint. Configure access separately.
Persistence, graphs and deployment
| Need | Pydantic AI ecosystem | Mastra ecosystem | Application still owns |
|---|---|---|---|
| Explicit branching | Pydantic Graph supports typed state, nodes/steps and edges | Workflow steps and control flow | Correct transitions and loop limits |
| Survive a process failure | Integrations with durable execution engines | Configured workflow persistence and recovery | External-effect deduplication and reconciliation |
| Evaluate behavior | Pydantic Evals | Agent/workflow evaluation tooling | Representative cases and acceptance criteria |
| Inspect execution | Instrumentation and observability integrations | Studio and tracing integrations | Redaction, retention and access |
| Deploy | Host the Python application and its dependencies | Managed or self-hosted deployment options | Capacity, networking, storage and incident response |
Pydantic Graph supports explicit graph construction; Pydantic AI is not limited to a single imperative loop. Its durable execution integrations include engines such as Temporal and DBOS. Stored conversation history and a recoverable active workflow solve different problems.
Mastra's deployment documentation describes several hosting paths and optional deployers. Verify runtime, storage adapter, streaming and job-duration compatibility. “Deploys to serverless” does not mean every long-running workflow can remain inside one HTTP request.
The current Mastra license mapping assigns Apache 2.0 to the core/general code and separate enterprise terms to designated directories. Do not describe the whole repository as Elastic License v2 or assume every enterprise feature has the core's terms.
Compare with LangGraph by testing a real workflow
LangGraph provides explicit state transitions and checkpoint-based execution. It can be used without adopting every LangChain component. Mastra also has workflows, and Pydantic has graph and durable-engine options. “Tool list versus state machine” is an incomplete selection rule.
Evaluate the same study-plan task with these questions:
- Can the team express the schema and domain rules clearly?
- Can it test the model boundary with controlled responses?
- What persists after an interruption, and where is it stored?
- What happens if saving succeeds but the response is lost?
- Can traces and plan data be exported without proprietary UI dependence?
- How much framework-specific code is required for the next likely feature?
A Python service may favor Pydantic AI; a TypeScript product may favor Mastra. An existing LangGraph system may be cheaper to extend than to replace. Crossing a language boundary has a cost, but can be reasonable when an independently operated service already exists. Benchmark cold starts and throughput rather than declaring one framework categorically faster.
Failure repairs and cost/benefit
| Problem | Repair | Tradeoff |
|---|---|---|
| Valid JSON names an inaccessible topic | Validate against server-derived accessible IDs | Extra catalog/access lookup |
| Parallel requests share mutable user dependencies | Per-request context; avoid mutable global identity | More explicit construction |
| Invalid output causes repeated model calls | Bounded repair, error classification and explicit fallback | Some requests fail instead of retrying forever |
| Resuming a run repeats a save | Idempotent operation ID and conditional version check | Persistent operation bookkeeping |
| Trace includes quiz history and credentials | Allowlist/redact before export | Less raw context for debugging |
| Provider swap changes tool behavior | Contract tests and representative evaluations | Migration takes more than changing a name |
For an illustrative workload of 10,000 plans, if 20% need one extra generation, there are 12,000 generations. Reducing that repair rate to 5% gives 10,500 generations, a 12.5% reduction in generation count. This is not automatically a 12.5% reduction in the entire service bill; model outputs, tokens, storage and fixed costs can differ.
Interview questions and answer checks
- Does
minutes: intprove the study plan fits a 90-minute budget? No. Validate individual values and the sum against the trusted budget. - Can a correctly typed tool leak another learner's quiz results? Yes, if authorization is absent or the account scope is model-controlled.
- Why use dependency injection? It separates trusted services/context from model arguments and makes service behavior substitutable in tests.
- Does schema validation prevent hallucination? It constrains structure; plausible but false values require evidence and domain checks.
- When do you need durable execution? When active work must survive failures or long waits. Ordinary saved chat history alone is different.
- How do you test this planner? Schema/domain unit tests, unauthorized-access tests, controlled model-output cases, and held-out educational-quality evaluations.
- Can a model provider be swapped without redesign? Sometimes, but verify output modes, tool semantics, streaming, quality, limits and cost first.
- What would your closing recommendation be? Choose the smallest framework that makes the required contracts clear and recovery demonstrable within the team's stack; keep policy and business state in application-owned boundaries.
Final notes
Remember shape → meaning → permission → effect. A schema checks shape. Domain rules check meaning. Authorization checks permission. Persistence and idempotency control the effect. Typed frameworks help organize these layers; the application must still connect and verify them.