Semantic Caching
Semantic caching reuses a stored model response when a new request is judged meaningfully equivalent to a previous one. It typically retrieves candidate requests through embeddings and applies a similarity rule, allowing reuse beyond exact text matches while requiring safeguards for freshness, context and user-specific information.
What it is
The cache stores a request representation and its response. For a new request, an embedding search finds nearby records; a similarity evaluator decides whether a candidate is suitable for reuse. This differs from exact-key response caching and from prompt or prefix caching, which reuses intermediate computation while still generating a response. Semantic proximity is only a proxy for equivalent answer requirements. Two requests can share vocabulary but differ in date, authorization or a crucial condition. Cache identity must therefore include relevant application context, and invalidation must reflect source or policy changes.
What the work involves
The practitioner defines eligible request classes, partitions records by relevant user and application context and evaluates matching thresholds. Freshness and invalidation rules should be explicit. Useful outputs include a cache policy and a labeled set of valid and invalid reuse pairs. Testing inspects false hits as well as missed opportunities, and measures embedding and lookup overhead. Dynamic or action-bearing requests may need bypass rules. Stored responses should retain enough provenance to determine whether the evidence and conditions that supported them remain applicable.
Illustrative example
A public handbook assistant receives differently worded questions about the same stable leave policy. The cache retrieves a prior request and reuses its answer only when the policy version and employee category match. A question about a different country's policy is a deliberately similar negative test. After the handbook changes, the old records are invalidated. The test checks that semantic closeness does not override the category or version requirements.
Limits and common mistakes
False matches can return an answer that is fluent but wrong for the current user or condition. Time-sensitive information and hidden context make reuse difficult. Similarity thresholds vary by embedding model and workload, so a generic value is unreliable. Quality requires strict context boundaries, tested negative pairs and effective invalidation. A high hit rate alone is not success: the cache must preserve correctness and avoid disclosing another user's data while producing a net resource benefit.
Prerequisites
- hardEmbedding Models
Similarity is computed over embeddings.
- mediumVector Databases
Cached queries are indexed for nearest-neighbor lookup.
Sources and further reading
- GPTCache documentation
Describes embedding retrieval, cached response storage and similarity evaluation for semantic response reuse.
Last updated: 2026-10-10