DSPy is a Python framework for composing language-model programs and optimizing parts of those programs against a chosen metric. You specify inputs, outputs, and program structure. An optimizer can search for better instructions or demonstrations using examples and evaluation feedback. Some optimizers support weight training; ordinary prompt optimization does not change model weights.
DSPy still creates prompts and makes model calls. Its value is a repeatable way to build and improve a program, not a guarantee of correct answers. This chapter focuses on implementation and release decisions. The prompt-optimization lesson develops the evaluation and search concepts in more detail.
Learn the six components
| Component | Responsibility | What it does not establish |
|---|---|---|
| Signature | Describe named inputs, outputs, types, and the task | Whether a generated claim is true |
| Module | Implement a reusable prediction or composed program | Authorization to perform external actions |
| Adapter | Convert signatures, examples, and inputs into model interactions and parse results | Identical behavior across providers |
| Example | Supply task inputs and, when available, reference outputs | That the labels represent future traffic |
| Metric | Score a prediction according to a defined objective | That the objective captures every product requirement |
| Optimizer | Search for a program configuration with a better measured score | A global optimum or guaranteed production improvement |
An optimizer's output is an application artifact: code/configuration plus selected instructions and examples, depending on the optimizer. It is not equivalent to a compiler proving program correctness. See the official signature guide and adapter guide.
Interview exercise: extract requirements from a design answer
Suppose a practice tool identifies requirements explicitly stated in a candidate's answer. It should distinguish “the system must support 10,000 requests per second” from a suggestion that this might be needed. The product must not silently invent a requirement and then grade the candidate against it.
Functional requirements
- Extract stated functional requirements and non-functional constraints.
- Return the exact supporting text for each extracted item.
- Return empty lists when the answer does not state requirements.
- Let the learner correct an extraction before using it in feedback.
- Record the program and rubric versions that produced the result.
Non-functional requirements
- Measure missed requirements and unsupported extractions separately.
- Enforce an output schema and bounded processing time.
- Keep private answers out of unrelated optimization datasets.
- Compare performance by topic and answer length.
- Retain a reproducible baseline and a rollback artifact.
Start with a single Predict module. Do not add a reasoning module, search optimizer, or repair loop until an observed error motivates it.
Define a precise interface
The following signature describes an extraction task. It assumes DSPy is installed; running a prediction additionally requires configuring an appropriate model. It is an interface example, not a complete grading service.
from typing import Literal
import dspy
from pydantic import BaseModel
class Requirement(BaseModel):
kind: Literal["functional", "non_functional"]
statement: str
evidence_quote: str
class ExtractRequirements(dspy.Signature):
"""Extract only requirements explicitly stated in the answer.
Quote the supporting text; do not add recommended requirements.
"""
answer: str = dspy.InputField()
requirements: list[Requirement] = dspy.OutputField()
extractor = dspy.Predict(ExtractRequirements)
Typed fields help describe and parse the expected output. Successful parsing is only the first check. A valid Requirement can still misclassify a suggestion or contain an invented quote. Independently check that each quote occurs in the supplied answer, then evaluate whether it supports the extracted meaning.
Use descriptive field names. evidence_quote is more informative than output_2. Avoid requesting a hidden internal reasoning transcript. If the product needs an explanation, request a brief evidence-based justification that can itself be reviewed.
Separate development from serving
Read diagram source
flowchart TD
A[Labeled development examples] --> B[Baseline program and metric]
B --> C[Bounded optimization search]
C --> D[Candidate artifact]
D --> E[Held-out tests and error review]
E --> F{Release criteria met?}
F -->|No| G[Retain baseline and inspect failures]
F -->|Yes| H[Versioned deployment]
U[New learner answer] --> H
H --> I[Prediction and independent validation]
I --> J{Accepted?}
J -->|Yes| K[Show editable extraction]
J -->|No| L[Bounded repair or clear failure]
Optimization belongs in a controlled development process. A user request normally invokes the saved program; it should not trigger a fresh expensive search over the training set.
Choose an optimizer deliberately
| Approach | What changes | Useful when | Main expense or risk |
|---|---|---|---|
| Manual baseline | Instructions and examples chosen by the developer | Establishing an understandable reference | Limited search coverage |
| BootstrapFewShot | Demonstrations selected from suitable traces/examples | Good examples are likely to improve behavior | Bad labels or a weak metric select bad demonstrations |
| MIPROv2 | Instructions and optionally demonstrations | Several interacting prompt choices need evaluation | Search calls and validation overfitting |
| GEPA | Prompt candidates informed by reflective feedback | Failure traces provide useful improvement signals | Feedback quality and search expenditure |
| Fine-tuning optimizer | Model parameters, where supported | Adequate data and a supported training path exist | Training, deployment, and model-specific constraints |
MIPROv2 bootstraps demonstration candidates, proposes instructions, and searches combinations with Bayesian optimization. It can also optimize instructions without demonstrations. There is no universal “10–20 prompts” or “100–500 calls” budget; actual work depends on candidates, examples, trials, program depth, and settings.
Do not compare optimizers using their own best development scores alone. Give them a comparable budget and evaluate the selected artifacts on untouched examples. Review rare but costly failures separately from the average.
Design a metric that cannot win by doing nothing
For the extraction exercise, define the matching rules between a predicted requirement and an annotated requirement before scoring. Review ambiguous labels with more than one annotator.
| Measurement | Example failure it exposes |
|---|---|
| Requirement recall | Omitting a stated latency target |
| Requirement precision | Inventing an availability target |
| Evidence support | Quoting text that does not justify the extraction |
| Empty-answer behavior | Hallucinating requirements when none are present |
| Per-topic results | Doing well on chat systems but poorly on payments |
| Latency and total model cost | Buying a small quality gain with excessive retries |
An “all quotes are substrings” metric alone can be maximized by returning no requirements. An exact-text metric can unfairly reject valid normalization. Combine task-aware metrics with hard acceptance rules and inspected examples. Keep near-duplicate answers and variants of the same exercise in the same data split to reduce leakage.
A numerical example: suppose a search tests 40 candidates on 60 examples, with two model calls per program execution. That is 40 × 60 × 2 = 4,800 task-model calls before proposal, feedback, retries, and final evaluation. At an assumed average $0.001 per task call, that component costs $4.80. These are illustrative assumptions, not a DSPy quote or default.
Runtime refinement is not a hard guarantee
Older tutorials describe DSPy assertions as the current way to enforce constraints. The official legacy assertions page marks that mechanism deprecated and unsupported. Current code should use supported APIs and explicit application validation.
dspy.Refine runs a module up to a configured number of attempts, uses a reward function and feedback, and selects a result. Crucially, the selected result may be the best available prediction without meeting the desired threshold. Feedback generation can add calls beyond the number of module attempts. See the current Refine API and implementation.
For the extraction task:
- Validate the response structure and evidence quotes.
- If a repair is useful and the shared time/cost budget allows it, attempt a bounded repair.
- Validate again outside the model's instructions.
- If the result still fails, display a clear extraction failure or request manual correction.
- Do not pass a rejected extraction into the grading stage as if it were trusted.
A prompt saying “do not reveal personal data” cannot establish a hard privacy property. Control what data is provided, where calls execute, and what outputs may be released. DSPy optimization and runtime refinement do not replace those controls.
Model changes and release management
A stable Python signature reduces some integration work when changing models. It does not ensure that every provider supports the same structured-output, tool-use, context, or generation behavior. An adapter may also use different formatting or fallback paths. Inspect the actual calls and failure modes.
Use this migration sequence:
- Record the existing model, dependency versions, adapter, instructions, demonstrations, and metric version.
- Run the current artifact on the proposed model without optimization.
- Compare quality, cost, latency, and failure slices with the incumbent.
- Optimize only if the expected improvement justifies the cost.
- Run untouched release tests and a limited rollout.
- Roll back on defined regressions and investigate them.
A new model release does not imply every prompt breaks. Recompilation does not automatically recover lost quality. Also inspect saved examples: an optimized prompt can embed training material, so the artifact needs the same data review as other application inputs.
Interview questions and answer notes
- Does DSPy remove prompts? No. It provides program abstractions and methods to optimize how model calls are prompted.
- Does a valid output type establish a correct extraction? No. Validate source support and task meaning separately.
- Why can optimization improve validation scores but hurt production? Metric mismatch, repeated selection on a small validation set, leakage, or a traffic shift.
- Does
Refine(N=3)mean exactly three billable model calls? No. A module can make several calls, feedback adds work, and early stopping can reduce attempts. - What happens when no refinement meets the threshold? The application must verify the selected result and apply its own failure policy; selecting the best result is not acceptance.
- Would you re-optimize immediately after changing providers? First evaluate the unchanged artifact and confirm integration behavior; then make a measured decision.
- Why save more than the optimized instructions? Reproduction also depends on code, examples, model/configuration, dependencies, adapter, and evaluation versions.
Final notes
Remember signature → program → metric → search → independent evaluation → release. Start with a baseline, make the metric hard to exploit, and separate a higher score from a production guarantee. The strongest interview answer explains which failures are caught by ordinary code and which require empirical evaluation.