Prompt Caching
Prompt caching reuses computation for a repeated input prefix so that later model requests need less repeated processing. It is an inference optimization for shared instructions or context; it differs from returning a previously generated answer because the model can still produce a new response to each request.
What it is
Transformer inference computes internal representations for input tokens before generating an answer. A serving system can retain eligible prefix computation and reuse it when another request begins with the same supported prefix. Cache matching, lifetime, minimum length and billing rules depend on the provider or engine. Stable instructions and reference documents can form a reusable prefix, while a changing question belongs after it. Prompt caching is distinct from semantic caching, which retrieves an old response for a meaningfully similar request, and from application memoization of an exact result.
What the work involves
The practitioner organizes messages so stable material precedes variable material, then measures actual cache reads and writes through provider usage fields or engine metrics. A useful experiment compares identical workloads with and without eligible prefix reuse, tracking latency and total cost rather than assuming every request hits. Cache policy also needs a tenant boundary and a plan for changed instructions or documents. The artifact is a request layout and measurement report that explain which portion was reused under the selected implementation.
Illustrative example
A support application repeatedly sends the same product manual and response policy, followed by a different customer question. The stable manual and policy are placed before the variable conversation. Later requests may reuse their prefix computation, while the assistant still reasons over the new question and generates a fresh answer. When the manual changes, the application uses the updated version and observes a new cache population rather than expecting the previous cached prefix to contain new information.
Limits and common mistakes
A cache hit is not a quality improvement, and a miss does not mean the application is broken. Small changes near the beginning can invalidate reuse; low request frequency can make writes or expired entries uneconomic. Provider caching semantics and pricing can change, so measurements need the actual service configuration. Reusing prefix computation also does not exempt an application from access controls or retention requirements for the material included in requests.
Prerequisites
- mediumContext Engineering
Caching is one technique within context engineering — understanding what to cache requires understanding context strategy
Related skills
- → is part of: Token Optimization
- → is part of: Context Engineering
- → is part of: Inference Optimization
Sources and further reading
- Prompt caching
Documents prefix reuse, cache configuration and usage accounting for the Claude API.
- Prompt caching
Documents provider-specific prefix matching and cached-token reporting.
Last updated: 2026-10-10