Speculative Decoding
Speculative decoding accelerates autoregressive generation by proposing several tokens with a cheaper draft process and checking them with the target model. A verification procedure accepts valid proposals and corrects rejected ones, reducing sequential target-model work when the draft agrees often enough to justify its overhead.
What it is
A draft model or another proposal mechanism produces candidate continuation tokens. The target model evaluates that continuation in parallel, and an acceptance-and-correction rule determines how much can be retained. Exact speculative sampling methods preserve the target distribution under their assumptions rather than simply accepting any fluent draft. This differs from using a smaller model as the final generator and from ordinary parallel request batching. Performance depends on proposal cost, agreement and the target model's execution characteristics. The algorithm changes the schedule of generation work, while the verification rule is what preserves the intended output distribution.
What the work involves
The practitioner selects a compatible proposal mechanism, configures draft length and measures acceptance behavior on real tasks. Benchmarks keep target model, sampling policy and hardware comparable. Useful outputs include decoding configuration and latency comparisons that account for the draft's resources. Testing covers prompts with low draft agreement, long outputs and concurrent requests. The implementation's exactness claims should be checked against the chosen verification and sampling options, especially when a runtime uses an approximate variant or additional heuristics.
Illustrative example
A service generates repetitive structured descriptions using a large target model. A smaller draft model proposes short token blocks, and the target verifies them. The team compares the same request set with and without speculation, including unusual terminology where acceptance is lower. It inspects whether the extra draft work still helps under concurrent load. A faster result on the easy subset does not justify enabling the method until the broader latency and output contract are checked.
Limits and common mistakes
Poor agreement can make drafting overhead outweigh saved target steps. Extra memory and compute may reduce serving capacity. Exactness belongs to a specified algorithm and implementation, not every feature called speculative decoding. Quality requires both distribution-preservation checks where claimed and realistic performance measurements. A technique that improves single-request decoding may behave differently under heavy batching, so throughput and interactive latency should be evaluated together for the deployment's workload.
Prerequisites
Speculative decoding uses a draft model to predict tokens that the main model verifies — understanding autoregressive generation is essential
- mediumLLM Inference Serving
Speculative decoding is implemented within inference engines like vLLM — practical experience with the engine helps
Related skills
- → is part of: Inference Optimization
Sources and further reading
- Fast Inference from Transformers via Speculative Decoding
Defines draft generation, target verification and sampling correction that preserve the target distribution under the method's assumptions.
Last updated: 2026-10-10