DSPy
DSPy is a framework for building language-model programs as composable modules with declared input and output roles, then optimizing their instructions or examples against a metric. It makes prompt behavior part of a program and evaluation workflow, while leaving the developer responsible for the task specification and evidence of improvement.
What it is
A DSPy program expresses tasks through signatures and modules rather than scattering hand-written prompts throughout application code. A signature describes what a call should consume and produce; modules can implement prediction, reasoning or retrieval-related steps. An optimizer evaluates the program on examples and adjusts supported prompt components or other configurable elements. The original DSPy work describes this as compilation, but it is not conventional compilation into machine instructions. The result is an executable language-model pipeline whose performance depends on the selected model, modules, data and metric.
What the work involves
The practitioner creates meaningful signatures, assembles the smallest program that performs the task and defines an evaluator. Training and validation examples are kept separate from the final test. Optimization is run with a recorded budget, then the selected program is inspected for invalid assumptions and evaluated end to end. A useful artifact includes code, model configuration, optimized state and test results. Retrieval failures or unreliable tool execution still need direct debugging; an optimizer cannot compensate for information the program never receives.
Illustrative example
A question-answering application separates query generation, document retrieval and answer production into modules. Its metric checks whether the answer is supported by the retrieved passage and whether it addresses the question. Optimization selects demonstrations for the query and answer modules. A held-out test then checks whether unfamiliar questions retrieve the right documents. The resulting program can be compared with the original hand-written pipeline under the same retrieval index and model settings, making the source of an improvement easier to investigate.
Limits and common mistakes
DSPy does not make evaluation objective by itself. A weak metric can reward fluent unsupported answers, and optimization on a small example set can produce brittle prompts. Library APIs and supported optimizers evolve, so reproducibility requires pinned versions and saved configuration. Claims about optimized performance should name the workload and budget. The framework is useful for organizing and improving programs, rather than a guarantee that programmatic prompting outperforms a simpler implementation.
Prerequisites
DSPy compiles and optimizes prompts programmatically — you must understand what manual prompt engineering does before automating it
- mediumModel Evaluation
DSPy optimizes prompts against metrics — you need to understand what metrics mean to define optimization targets
Related skills
- → is an instance of: Automated Prompt Optimization
Sources and further reading
- DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
Primary explanation of signatures, modular programs and prompt optimization.
- DSPy repository
Maintained source and documentation for the framework's current implementation.
Last updated: 2026-10-10