DSPy is a framework for expressing language-model programs and optimizing them against examples and a metric. It separates the program's input/output contracts and control flow from some of the instructions and demonstrations used to implement them. An optimizer searches for a better configuration; it does not prove that it has found the best possible prompt.
The original DSPy paper describes compiling declarative language-model calls into improved pipelines. The framework is one option for systematic optimization, not a requirement for every LLM application. Khattab et al..
Learn the four building blocks
| Element | Meaning | Example |
|---|---|---|
| Signature | Declares task inputs and outputs | Ticket text → routing label |
| Module | Implements one or more model calls and program behavior | A classifier followed by evidence validation |
| Metric | Scores whether the program achieved the task | Correct label and a valid source excerpt |
| Optimizer | Searches instructions, examples or other supported parameters | Select demonstrations using development feedback |
A signature describes a contract. It does not supply retrieval, authorization or a correct algorithm merely because its description says “multi-hop reasoning.” If the task requires several searches, the program must actually perform them.
Start with a program that can be evaluated
For the support-routing task, define:
- Inputs: ticket text and the versioned routing rubric.
- Outputs: one allowed label and an exact supporting excerpt, when present.
- Control flow: predict, validate the output contract, then return a routing proposal.
- Evaluation: correct label, excerpt grounded in the ticket, and explicit handling of ambiguous cases.
- Boundaries: no account modifications or private cross-tenant examples.
The core task contract can be expressed without any framework:
route(ticket_text, rubric_version) -> label, evidence_excerpt
Allowed labels: cancel, refund, multiple, other
Evidence excerpt: exact substring of the input, or empty when appropriate
Side effects: none
Implement this as a minimal baseline before adding an optimizer. A framework cannot make an undefined label policy measurable.
Choose an optimizer for the actual problem
| Optimizer family | What it changes | What to inspect |
|---|---|---|
| Labeled or bootstrapped few-shot | Demonstrations; bootstrapping can use successful program traces | Example correctness and representativeness |
| MIPROv2 | Instructions and demonstrations, using candidate evaluation and search | Search budget, development-set overfitting |
| GEPA | Prompt candidates informed by reflection on trajectories and feedback | Quality and privacy of feedback; objective gaming |
| BootstrapFinetune | Model weights using generated/selected training traces | Training quality, supported model and deployment costs |
These distinctions reflect the current DSPy optimizer documentation. Prompt optimization usually leaves the base model weights unchanged. Some DSPy optimizers explicitly perform fine-tuning, so “DSPy never changes weights” would also be wrong.
“Compile” in this context means producing an optimized program configuration. It does not imply a formal correctness proof or necessarily gradient descent. Avoid assuming that a prompt is literally a differentiable neural-network weight.
Keep evaluation independent of selection
Read diagram source
flowchart LR
T[Training examples and traces] --> O[Generate candidate programs]
D[Development cases and metric] --> O
O --> S[Select and freeze candidate]
S --> H[Evaluate on untouched test set]
H --> R[Release or investigate failures]
Optimizer APIs differ in how they use training and validation inputs. Whatever the interface, keep a final test set outside the search. Repeatedly choosing prompts based on that test set turns it into another development set.
Split related examples by the relevant unit: customer, document family, time period or task instance. Randomly splitting near-duplicate tickets can exaggerate generalization. Review rare labels, adversarial instructions and inputs requiring abstention separately.
Design a metric the system cannot cheaply exploit
Exact match is useful for a fixed routing label. It is less appropriate for an open-ended answer with several correct phrasings. A model judge can assess a rubric, but needs calibration against human-reviewed cases and checks for position, style and length bias.
Use hard acceptance conditions for properties that cannot be traded away. An aggregate score should not allow a modest accuracy improvement to compensate for unauthorized data access. Evaluate those boundaries independently of the prompt optimizer.
For the classifier, report label accuracy, per-label errors, evidence validity, invalid-output rate and latency. If the objective rewards only accuracy on common cases, the optimizer may find a short prompt that performs badly on rare categories.
Budget the search and the deployed program
An illustrative search with 100 candidate evaluations, 50 cases per evaluation and two model calls per case uses 100 × 50 × 2 = 10,000 task-model calls. Teacher generation, reflection, retries and final evaluation add work. Some algorithms reuse results or evaluate subsets; inspect actual accounting rather than assuming this full schedule.
There are two costs to justify: optimization cost and steady-state serving cost. An optimized program with longer demonstrations may improve quality while increasing every future request's input tokens. Compare total lifecycle cost and the cost per successful task.
Release and model changes
Version the program, optimized instructions/examples, model, adapters, generation settings, dataset and metric. Preserve the baseline and rollback path. Monitor real traffic for distribution changes and investigate new failures before searching again.
For a model upgrade, evaluate the existing program first. Re-optimize if the results justify it. A model change does not logically require rewriting every prompt, and recompilation does not automatically restore previous quality.
Interview practice
Q1: What problem does DSPy solve?
It provides structure for composing model programs and searching configurations against a metric. It reduces ad hoc prompt selection when there are representative examples and a meaningful evaluation. It does not remove the need to define the task, data access or acceptance criteria.
Q2: Does a signature implement a multi-hop retrieval system?
No. It describes what a call receives and returns. The program must implement retrieval, intermediate state and stopping behavior. A descriptive class name cannot create missing control flow.
Q3: How does MIPROv2 differ from GEPA at a high level?
MIPROv2 searches instruction and demonstration candidates using performance feedback. GEPA uses reflective feedback on execution trajectories to propose prompt changes. I would select based on the available metric/feedback, search budget and measured results rather than declaring one universally superior.
Q4: Is prompt optimization fine-tuning?
Not when it changes instructions and examples only. Fine-tuning changes model parameters. DSPy supports both kinds of optimization through different components, so I would specify exactly what the chosen optimizer changes.
Q5: Why might the highest-scoring candidate be a poor release?
It may overfit the development cases, exploit a weak judge, increase serving cost or fail an important category. Freeze the candidate, use an untouched test set, inspect failure slices and apply separate operational and security acceptance criteria.
Q6: What if there are only thirty reviewed examples?
Start with a simple baseline and a modest search. Preserve independent checks, inspect every failure and collect more representative data. A large automated search against a tiny dataset can select noise. More search is not a substitute for better evidence.
Final notes
Recall card: Contract → program → metric → search → independent test. Automated prompt optimization is an empirical development process, with versioning and release controls like the rest of the application.
Related: few-shot learning, structured generation, fine-tuning.