Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Building tool-use agents: contracts, execution, and evidence

By Anup Rai16 min readReviewed September 2026

A tool-use agent combines model-directed decisions with application-controlled operations. The model selects a tool and proposes arguments. Your application decides whether those arguments are valid and authorized, performs the operation, and returns its actual outcome.

A tool contract defines inputs, outputs, side effects, errors, and relevant execution guarantees. A JSON schema describes part of that contract. It does not establish identity, authorization, factual correctness, or transaction safety.

Practice objective: build a support assistant that finds the correct customer, prepares a ticket, creates it when authorized, and reports the saved ticket ID. The difficult cases are two customers with similar names, a changed permission, a retried write, and a backend timeout after creation. Solve those cases before adding dozens of tools.

Current protocol and SDK references were checked on September 24, 2026. Prerequisites: tool-agent architecture, MCP foundations, and durable execution.

Start with the business operation

Tool Purpose Important boundary
search_customers Return a small authorized candidate set A search match is not proof of intended identity
get_customer Read an identified authorized record Validate current object access on every request
prepare_ticket Validate a proposed ticket and produce a reviewable intent Preparation does not authorize creation
commit_ticket Create the exact authorized intent Deduplication, current policy, and transaction checks
get_ticket_operation Inspect the outcome of an earlier attempt An unknown outcome must remain visible

Separate operations when they have different permissions, side effects, or review needs. A composite operation is useful when it expresses a stable business capability with clear semantics. “Every tool must do exactly one low-level step” is too rigid; create_ticket_with_initial_note can be a valid atomic domain operation.

A complete contract answers these questions

  1. Selection: when should the agent use this operation, and which similar operation should it prefer for another need?
  2. Arguments: which types, formats, limits, defaults, and units are accepted?
  3. Authority: which identity, tenant, object permission, and approval are required?
  4. Result: which fields establish success, absence, partial data, or pending work?
  5. Effects: what changes, and can repetition create another effect?
  6. Failure: which errors permit correction, retry, reconciliation, or escalation?
  7. Resources: what deadlines, pagination, concurrency, and spending limits apply?

Use meaningful names, concise descriptions, and examples of ambiguous cases. A useful description improves selection; it does not enforce policy.

Example input schema

This is an application/MCP input schema for an authorized customer search. The tenant comes from trusted request context, so it is absent from model arguments.

{
  "type": "object",
  "properties": {
    "query": {
      "type": "string",
      "minLength": 2,
      "maxLength": 120,
      "description": "Customer name, email, or known customer ID. This searches individual customers, not aggregate counts."
    },
    "limit": {
      "type": "integer",
      "minimum": 1,
      "maximum": 10,
      "default": 5
    }
  },
  "required": ["query"],
  "additionalProperties": false
}

default is an annotation: JSON Schema validation does not itself insert the missing value. Apply an explicit application default or use a validator that deliberately supports default insertion. See JSON Schema annotations.

Strict generation is useful, but has a defined scope

Claude's strict: true constrains generated tool arguments to its supported schema grammar. It does not prove that a customer exists or that a write is permitted. The supported subset also matters: current structured-output guidance excludes numeric bounds and some length constraints from the generation grammar; SDK transformations can simplify the sent schema and validate the original constraints afterward. See strict tool use and schema limitations.

Keep the complete schema at the server boundary. Use a provider adapter to produce a supported model-facing schema. Test both. Never remove business validation because a model or SDK returned valid JSON.

Validate in layers

Architecture / visual model
flowchart TD A[Proposed arguments] --> S[Parse and validate full schema] S --> I[Attach trusted request identity] I --> Z[Authorize operation and target objects] Z --> B[Check business state and approval] B --> R[Reserve operation and execution budget] R --> X[Execute bounded transaction or service call] X --> V[Validate and minimize returned result] V --> O[Record outcome and return evidence]
Read diagram source
flowchart TD
    A[Proposed arguments] --> S[Parse and validate full schema]
    S --> I[Attach trusted request identity]
    I --> Z[Authorize operation and target objects]
    Z --> B[Check business state and approval]
    B --> R[Reserve operation and execution budget]
    R --> X[Execute bounded transaction or service call]
    X --> V[Validate and minimize returned result]
    V --> O[Record outcome and return evidence]
Layer Example rejection What it protects
Syntax/schema limit: true, extra tenant_id, overlong subject Interface integrity
Identity Missing or expired authenticated session Attribution
Object authorization Customer belongs to another tenant Data and action scope
Business validation Closed customer cannot receive the requested ticket type Domain invariants
Approval/precondition Prepared intent changed after review Authorized action fidelity
Execution Unique operation already committed Duplicate prevention
Result validation Backend returned an incompatible or incomplete record Truthful downstream interpretation

Executable exercise: construct a scoped ticket command

This original helper validates a proposal and attaches a server-owned tenant. It does not insert a ticket. authorized_customer_ids is a trusted permission snapshot supplied by the application, never an argument taken from the model.

def prepare_ticket_input(raw, *, tenant_id, authorized_customer_ids):
    import unicodedata

    if type(raw) is not dict:
        raise ValueError("ticket arguments must be an object")
    if set(raw) != {"customer_id", "subject", "priority"}:
        raise ValueError("provide customer_id, subject and priority only")
    if type(tenant_id) is not str or not tenant_id.strip():
        raise ValueError("trusted tenant is required")
    if type(authorized_customer_ids) not in (set, frozenset):
        raise ValueError("trusted authorization set is required")
    if any(type(item) is not str or not item for item in authorized_customer_ids):
        raise ValueError("invalid authorization set")

    customer = raw["customer_id"]
    subject = raw["subject"]
    priority = raw["priority"]
    if type(customer) is not str or not 1 <= len(customer) <= 80:
        raise ValueError("invalid customer reference")
    if type(subject) is not str or not 1 <= len(subject.strip()) <= 160:
        raise ValueError("subject must contain 1 to 160 characters")
    if any(unicodedata.category(char) in {"Cc", "Zl", "Zp"} for char in subject):
        raise ValueError("subject must be a single line without control characters")
    subject = subject.strip()
    if type(priority) is not str or priority not in {"low", "normal", "high"}:
        raise ValueError("invalid priority")
    if customer not in authorized_customer_ids:
        raise PermissionError("customer is unavailable for this operation")
    return {"tenant_id": tenant_id, "customer_id": customer,
            "subject": subject, "priority": priority}

The write path must recheck current access and business state while committing, using an appropriate transaction or authorization revision. A previously valid permission snapshot is not a permanent grant. Also impose transport/body size limits before parsing; this function is not protection against arbitrarily large incoming requests.

Money and quantities require explicit units

A financial tool should not accept an unspecified floating-point amount and assume that positive means safe. Use a decimal or integer minor-unit representation with an explicit currency and documented scale, validate finite/range constraints, establish both accounts' authority, and delegate balance/ledger invariants to the transaction service. A $10,000 threshold by itself establishes none of those guarantees. See the payment decision case for a broader design discussion.

Return evidence the next step can use

An illustrative search result for a public training fixture is:

{
  "status": "complete",
  "customers": [
    { "customer_id": "cust_training_17", "display_name": "North Region Training Account" }
  ],
  "next_cursor": null,
  "observed_at": "2026-09-24T14:30:00Z"
}

Return only the fields needed for the task. Use stable opaque IDs obtained from authorized results for follow-up calls; do not force the model to invent database identifiers. Human-readable names are useful for search, but name ambiguity is a reason to clarify, not a reason to pick the first result.

For paginated search, distinguish returned count from total matches. Do not claim a total you did not compute. A cursor should preserve query/sort context and authorization checks; a cursor is not an access token.

Expose the service through MCP

MCP is a protocol between a host application and servers that expose tools, resources, and prompts. A compatible host translates discovered capabilities into a model's interface and executes calls. The model does not automatically speak every MCP feature, and connecting a server does not guarantee that every host supports it.

Architecture / visual model
flowchart LR M[Model provider] <--> H[Agent host and policy] H <--> C[MCP client adapter] C <-->|stdio or Streamable HTTP| S[MCP server adapter] S --> D[Domain service] D --> B[Database or external API] S --> A[Authentication and request scope] A --> D
Read diagram source
flowchart LR
    M[Model provider] <--> H[Agent host and policy]
    H <--> C[MCP client adapter]
    C <-->|stdio or Streamable HTTP| S[MCP server adapter]
    S --> D[Domain service]
    D --> B[Database or external API]
    S --> A[Authentication and request scope]
    A --> D

Current version boundary

The 2026-07-28 protocol carries version/client metadata on each request. It does not use the older initialization handshake. server/discover provides server/version information; clients may inspect it or send a request and handle an unsupported-version response. A dual-era SDK can also support older clients. See versioning and compatibility.

tools/list exposes available tool definitions; tools/call invokes them. The current result envelope distinguishes completion from input_required. Tool definitions can declare output schemas for structured results. Discovery may be paginated and authorization-dependent, so fetch the applicable pages and scope any cache correctly. See the tool specification.

Deployment Appropriate use Important operating detail
stdio child process Local host launches a tool service stdout belongs to protocol messages; send diagnostics elsewhere
Streamable HTTP A remote service shared by permitted clients Authenticate requests, enforce origin/host policy and rate limits
Both through adapters Local development and remote operation Keep the same domain contract; test each transport/version path

The MCP HTTP authorization specification covers access tokens, resource identity, metadata, and client registration. It does not replace per-customer authorization in your domain service. Do not forward an incoming bearer token to an unrelated downstream API. See MCP authorization.

Minimal TypeScript server: public practice catalog

This local teaching service returns three fixed topic identifiers and accesses no customer data. The current TypeScript SDK uses split v2 packages and Standard Schema-compatible libraries. See the official TypeScript SDK.

import { McpServer } from '@modelcontextprotocol/server';
import { StdioServerTransport } from '@modelcontextprotocol/server/stdio';
import * as z from 'zod/v4';

const server = new McpServer({ name: 'practice-catalog', version: '1.0.0' });
server.registerTool(
  'list_practice_topics',
  {
    description: 'List the fixed public topic IDs available in this practice fixture.',
    inputSchema: z.object({}),
    outputSchema: z.object({ topics: z.array(z.string()) }),
  },
  async () => {
    const result = { topics: ['tool-contracts', 'authorization', 'recovery'] };
    return {
      content: [{ type: 'text', text: JSON.stringify(result) }],
      structuredContent: result,
    };
  },
);
await server.connect(new StdioServerTransport());

Equivalent Python registration

The current official Python SDK v2 uses MCPServer. Its older mcp.server.fastmcp.FastMCP examples belong to the v1 line; the separate FastMCP project also has its own APIs. Pin the intended package/version rather than treating these as identical. See the official Python SDK.

from mcp.server import MCPServer

mcp = MCPServer("practice-catalog")

@mcp.tool()
def list_practice_topics() -> dict[str, list[str]]:
    """List the fixed public topic IDs available in this practice fixture."""
    return {"topics": ["tool-contracts", "authorization", "recovery"]}

With the v2 CLI extra installed, the SDK's mcp run server.py command runs a module like this; use its documented transport options. The return annotation produces structured output under the current SDK. See first steps and structured output.

Both snippets are reference-checked registration examples, not a production server deployment. Before exposing private tools, add authentication, request-scoped identity, domain policy, bounded I/O, structured results, and integration tests. SDK installation alone supplies none of those business decisions.

Make discovery useful without making it authoritative

Static registration is reasonable for a small, stable tool set. Dynamic discovery is useful when the catalog is large or permissions differ. Neither is universally mandatory.

A two-stage selection design can:

  1. Filter candidate tools by the caller's permitted domain and capabilities.
  2. Retrieve likely tools from names, descriptions, examples, and argument meaning.
  3. Load their full schemas into the model context.
  4. Let the model select an operation or request more discovery.
  5. Revalidate authorization at execution even when the catalog already filtered it.

Claude's tool-search interface supports deferred definitions. Current guidance requires sending full deferred definitions in the API request even though they are not immediately inserted into model context. The search tool itself remains available. This distinction matters when estimating context versus payload size. See tool search.

Worked example: 200 schemas averaging 250 tokens occupy about 50,000 tokens. Eight selected schemas occupy about 2,000 tokens before search metadata and history. The 96% schema-token reduction is arithmetic, not a guaranteed quality or billed-cost improvement. Discovery can miss the needed tool or add latency.

Measure retrieval recall, wrong-tool selection, extra discovery calls, task success, and total billed cost. Tune the selected count to the workload. There is no universal rule that accuracy collapses at 50 tools or that five tools are always enough.

Compose tools at the right level

Pattern Who controls intermediate steps? Benefit Cost or limitation
Model-directed sequence Model reasons after observations Handles ambiguity and changing goals Repeated inference and larger history
Model-generated program Code loops, branches, aggregates Reduces repeated model reasoning over routine data Code runtime, resource bounds, and policy enforcement
Fixed server composition Application-owned workflow Clear domain contract and efficient repeated execution Less flexible; composite failures must be explicit
Durable workflow Persisted state machine coordinates steps Recovery, waits, approvals, and compensation More state and operational complexity

Programmatic tool calling reduces model round trips, not all network requests. Client-hosted tools can still require calls and result continuations while code is paused. It also needs authorized tool execution and resource limits. Anthropic's current allowed_callers controls presentation/calling mode, not a hard security boundary. See programmatic calling.

Example: gather 20 independent read results at 200 ms each. Sequential I/O needs about four seconds. Concurrency limited to five gives four waves, about 0.8 seconds under ideal equal-duration conditions, plus scheduling/network overhead. Code can aggregate those results without asking the model to read every row. Actual backend quotas and tail latency limit the gain.

For customer search → ticket creation, code must not select results[0] merely because the result list is nonempty. Require an unambiguous authorized customer or an explicit selection, then prepare the exact ticket. Flexibility must not erase identity checks.

Make writes recoverable

Architecture / visual model
sequenceDiagram participant H as Agent host participant S as Ticket service participant D as Database H->>S: Prepare validated ticket under trusted identity S-->>H: Intent ID, payload hash, version and expiry H->>S: Commit authorized intent with stable operation ID S->>D: Begin transaction and lock applicable state S->>D: Recheck authorization, intent and business conditions S->>D: Claim operation, create ticket, save result and outbox D-->>S: Commit transaction Note over H,S: Reply may be lost after commit H->>S: Read outcome or repeat exact operation under contract S->>D: Read saved operation result S-->>H: Same ticket ID, or explicit unresolved/conflict outcome
Read diagram source
sequenceDiagram
    participant H as Agent host
    participant S as Ticket service
    participant D as Database
    H->>S: Prepare validated ticket under trusted identity
    S-->>H: Intent ID, payload hash, version and expiry
    H->>S: Commit authorized intent with stable operation ID
    S->>D: Begin transaction and lock applicable state
    S->>D: Recheck authorization, intent and business conditions
    S->>D: Claim operation, create ticket, save result and outbox
    D-->>S: Commit transaction
    Note over H,S: Reply may be lost after commit
    H->>S: Read outcome or repeat exact operation under contract
    S->>D: Read saved operation result
    S-->>H: Same ticket ID, or explicit unresolved/conflict outcome
Record Important fields Purpose
Prepared intent Tenant, actor, customer, normalized payload/hash, revision, expiry Binds the reviewable proposal
Authorization evidence Actor, scope, target, decision/revision, approval if required Establishes permitted action
Operation Scoped operation ID, payload hash, state, result ID Deduplicates exact requests
Ticket Tenant, customer, subject, priority, creation identity Authoritative business record
Outbox event Operation, event type, delivery state Publishes committed changes reliably

Within one database, enforce uniqueness and state transitions transactionally. Reusing the same operation ID with a different payload must be a conflict, not a new ticket or the old result disguised as success. Scope keys to the tenant and operation type and authenticate result lookup.

For a remote backend, your local transaction cannot atomically commit the remote effect. Use the backend's documented idempotency/reconciliation contract. A timeout becomes unknown when completion cannot be established. Retrying with a freshly generated key defeats deduplication. See execution and retry patterns.

Errors need an action, not just a sentence

Domain outcome Agent/controller response
Invalid arguments Correct the specific fields within bounds
Target absent or inaccessible Report limited information; do not reveal another tenant's records
Ambiguous target Ask for a discriminating field or explicit selection
Authorization denied Stop or use the defined authorization flow
Stale intent/conflict Refresh state and obtain applicable authorization for the changed action
Rate limited Honor a validated delay within the shared deadline/attempt budget
Unknown write result Reconcile before another potentially duplicating action
Partial read result Label missing coverage and avoid claiming completeness

These are application outcomes, not a universal list of MCP error codes. Preserve the actual provider/protocol error envelope and map it into a documented internal policy. Do not expose stack traces, credentials, or unauthorized identifiers as “helpful error context.”

Package procedures as skills

A skill is reusable procedural guidance with metadata and optional supporting files. The Agent Skills format requires a SKILL.md; scripts and references are optional. Its experimental allowed-tools support varies by host and is not a portable replacement for execution policy. See the Agent Skills specification.

An original practice skill might use this structure:

support-ticket-review/
  SKILL.md
  references/priority-policy.md
  scripts/validate_training_fixture.py

The skill should explain how to identify the correct customer, prepare a ticket, verify permitted scope, and report evidence. It should reference the current tool contract instead of embedding a stale copy of every schema. Installing a skill does not universally create MCP tools or insert everything into a system prompt; loading and tool registration depend on the host.

Reuse domain logic across HTTP and MCP

An HTTP framework can generate an OpenAPI description from typed request models. That document is not automatically a ready-to-use tool list for every model provider. An adapter must select operations, convert the supported schema subset, attach authenticated identity, execute the request, and return a bounded result.

Keep the domain service independent of transport. HTTP, MCP, and an administrative UI should reach the same authorization and transaction rules. Otherwise the assistant may find a path that bypasses checks enforced in the human interface.

Test behavior at four levels

  1. Pure contracts: wrong types, missing/extra fields, bounds, normalization, and result shapes.
  2. Service integration: actual database constraints, concurrent commits, authorization changes, and downstream timeouts.
  3. Protocol integration: discovery, pagination, version compatibility, cancellation, errors, structured results, and authentication.
  4. Agent evaluation: correct task completion, clarifications, abstentions, unsafe actions, tool choice, cost, and recovery.

Do not require a single exact tool sequence when several permitted sequences achieve the same outcome. Assert required invariants and forbidden effects, then use traces to diagnose deviations. A decline in tool selection can come from discovery, changed descriptions, schema conversion, context, or model behavior; it does not identify its own cause.

Evaluation case Expected outcome Forbidden outcome
Two authorized customers share a name Clarify identity Select the first result and write
Same operation submitted concurrently One business record under the contract Duplicate tickets
Same key with changed priority Conflict Silently return the first result as the changed action
Permission revoked after preparation Recheck and deny Use stale authorization to commit
Backend times out after commit Read/reconcile existing result Blind retry with a new key
Search returns partial pages Continue within budget or label partial Claim all records were examined
User asks a general explanation Answer without unnecessary private tools Query customer data without need

Start with representative task and failure slices, not a magic count of 100 questions. Expand the held-out set as actual errors appear. Use repeated trials and uncertainty estimates for model comparisons. See evaluation design.

Record enough to explain a result

Propagate a trace across model request, tool dispatch, domain service, and backend operation. Record tool/schema version, duration, outcome category, retry count, output size, and operation ID where appropriate. Raw arguments and outputs may contain private data; redact, sample, or omit them according to policy.

Metric Useful interpretation Common mistake
Verified task completion Intended result actually achieved Counting HTTP 200 as task success
Unauthorized/duplicate effects Critical correctness failures Hiding them inside an average success rate
Unknown outcomes and age Work needing reconciliation Treating timeout as known failure
p95/p99 tool/task latency User delay and deadline pressure Alerting only on average latency
Human minutes per task Operational burden Reporting automation rate without review effort
Cost per verified outcome Complete unit economics Counting only one model response

Set thresholds from the business SLO and observed baseline. One universal 95% tool-success threshold is inappropriate for both an optional search and a ticket-creation transaction. Avoid per-customer or per-request metric labels that create uncontrolled cardinality; use protected trace/log fields for those identifiers.

Version the behavior, not only the name

An optional input field can still change schema loading, strict-provider compatibility, defaults, or model selection behavior. An added output enum member can break a closed consumer. Treat “additive” as a compatibility hypothesis to test.

  1. Identify the changed input, output, side effect, error, permission, or timing contract.
  2. Maintain a contract version and test representative existing callers.
  3. Use a new tool/version when the old meaning cannot be preserved safely.
  4. Measure deprecated usage and provide a migration window and owner.
  5. Retire old behavior according to risk and policy; a security issue may require immediate disabling.

A deprecation description can help the model migrate, but cannot guarantee it will never call the old tool. Enforce lifecycle at the registry and service boundary.

Design review: flaws, repairs, and tradeoffs

Tempting shortcut Repair Benefit Cost
Expose the whole internal API Curate operations and permission scopes Smaller useful capability set Adapter maintenance
Trust strict JSON as authorization Validate identity, objects, and state Correct business boundary Extra policy and storage calls
Let the model pass tenant_id Derive it from verified request context Prevents accidental scope selection Request-context plumbing
Hide every workflow inside one giant tool Use clear domain operations and visible states Better review and recovery More explicit contracts
Add an idempotency key without atomic storage Enforce scoped uniqueness and payload binding Correct concurrent behavior Transaction/retention design
Log every argument and result Retain minimal protected evidence Better privacy and manageable volume More deliberate diagnostics

Interview exercise: support tools for a shared platform

Functional requirements

  1. Search authorized customers and resolve ambiguity.
  2. Prepare a ticket with validated subject and priority.
  3. Create the exact authorized ticket once under the service's deduplication contract.
  4. Retrieve task/operation outcome after failures.
  5. Expose compatible tools to selected local and remote hosts.
  6. Support audit review and tool-version migration.

Non-functional requirements

  1. Enforce tenant and object authorization at the service boundary.
  2. Bound latency, retries, result size, concurrency, and task cost.
  3. Preserve correct outcomes under concurrent requests and lost replies.
  4. Keep sensitive inputs and results out of unnecessary context and logs.
  5. Measure task correctness, recovery backlog, and user review effort.

Basic design: model → tool handler → ticket database. First flaw: arbitrary customer IDs are accepted. Add trusted identity and object checks. Second flaw: retries create duplicates. Add a prepared intent, scoped operation ledger, and transactional result. Third flaw: a remote service cannot share the transaction. Add explicit unknown state and reconciliation.

Sizing example: 3,000 tasks across eight hours, six tool calls/task, and a 5× sustained peak imply about 3.13 tool calls/second. At a 0.3-second mean service time, mean active tool work is about 0.94 requests. Two slots at 60% planned occupancy are an initial estimate, before backend limits, bursts, long calls, and separate model concurrency. More tools in the catalog do not inherently mean more concurrent requests.

Interview questions and answer notes

  1. What does a tool schema fail to specify? Identity, authorization, business truth, side effects, and recovery guarantees need additional contracts.
  2. Does JSON Schema default populate an omitted argument? Not by validation alone; apply an explicit defaulting policy.
  3. Why retain server validation with strict generation? Provider subsets, other callers, evolving state, and business rules remain outside that guarantee.
  4. Are opaque IDs bad tool inputs? No, if they come from an authorized result; they avoid ambiguous name resolution.
  5. Does MCP discovery grant execution permission? No. Execute only under current authenticated scope and domain policy.
  6. Is a skill the same as an MCP server? No. One packages guidance; the other exposes protocol capabilities.
  7. Does programmatic calling use only one network request? Not necessarily. It avoids repeated model reasoning but still performs tool I/O and possible result continuations.
  8. When is server-side composition appropriate? When a stable business operation can own its validation, transaction, and failure contract.
  9. What is wrong with generating a new retry key? It describes a new operation to the deduplication service and can duplicate the effect.
  10. Should an evaluation require one exact tool sequence? Only when the order itself is required; otherwise evaluate outcomes, constraints, and unnecessary work.
  11. Is an optional field always backward compatible? No. Test actual schemas, defaults, consumers, and model behavior.
  12. What would you emphasize in the closing answer? Clear contracts, trusted identity, bounded execution, recoverable writes, and measured task outcomes.

Final notes

Remember contract → scope → validate → execute → reconcile → evaluate. Start with a few useful operations that behave correctly under failure. Grow the catalog only when a new capability solves a concrete task and retains the same policy and evidence standards.

Next: Tool-agent use cases.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Computer-use agents: from screen observations to verified outcomes
NEXT LESSONTool-agent use cases: choose the workflow, prove the value →

Explore the diagram