Skip to main content
ExplainerSpeculative DecodingGoogle Research· 7 min read· in Artificial Intelligence

How Speculative Decoding Uses Residual Probabilities to Guarantee Lossless AI Generation

By pairing a fast draft model with a slower target model, speculative decoding accelerates text generation without altering the final output. When the target model rejects a drafted word, it samples from a residual probability distribution to perfectly correct the error.

By Mateo Ramos

In short

  1. Speculative decoding accelerates AI text generation by having a small model draft multiple words while a large model verifies them in parallel.
  2. The process is mathematically lossless, meaning the final output distribution perfectly matches what the large model would have produced alone.
  3. When a drafted word is rejected, the system samples from a residual probability distribution to correct the error and maintain accuracy.

A standard large language model generates text at roughly 40 tokens per second, which translates to a 25-millisecond wait for every single word. Speculative decoding can push that rate past 120 tokens per second on the exact same hardware, without altering a single comma of the final output.[2]

The mechanism behind this acceleration relies on pairing two distinct models: a massive, slow "target" model and a lightweight, fast "draft" model. The small model cheaply guesses the next several words in the sequence, and the large model verifies them simultaneously.

This parallel verification is the core of speculative decoding, a technique introduced by researchers Yaniv Leviathan, Matan Kalman, and Yossi Matias. Because a graphics processing unit is bottlenecked by loading weights into memory rather than performing the arithmetic, evaluating five drafted tokens takes roughly the same time as evaluating one.[1][2]

If the draft model guesses correctly, the system outputs multiple tokens in a single 25-millisecond forward pass. The speedup is profound, often reaching 2.5x to 3x in production environments, according to Google Research, fundamentally altering the latency profile of generative artificial intelligence.

By evaluating multiple tokens simultaneously, speculative decoding can triple effective throughput.

But the true breakthrough of speculative decoding is not the drafting itself; it is the mathematical guarantee that the output remains identical to what the slow model would have written alone. This property, known as lossless acceleration, ensures that speed never comes at the expense of reasoning quality.

The Mathematics of Acceptance

To maintain this identical output distribution, the system does not simply accept or reject words based on a flat confidence threshold. Instead, it meticulously compares the probability distributions of both models for every proposed token.[1]

Let the target model's probability for a specific word be $p(x)$, and the draft model's probability be $q(x)$. If the target model assigns a higher or equal probability to the word, the token is accepted immediately, as the draft model's guess aligns perfectly with the target's intent.[1]

The complexity arises when the draft model proposes a word that the target model considers less likely. The system does not automatically discard the word. Instead, it accepts the token with a probability equal to the ratio of the two distributions: $p(x)$ divided by $q(x)$.[1]

This ratio-based acceptance is a strict form of rejection sampling. It ensures that over thousands of generations, the frequency of any given word exactly matches the target model's original intent, even if the draft model over-proposed it during the parallel generation phase.[1]

"In conventional speculative decoding, the draft and target models are independent, and acceptance depends on how closely the draft distribution matches the target distribution," notes an analysis of the algorithm's performance by GMI Cloud.[2]

Tokens are accepted immediately if the target model agrees, but face rejection sampling if the draft model overestimates their likelihood.

Sampling the Residual Probability

When a drafted token is finally rejected, the system must recover gracefully. It discards the rejected word and all subsequent drafted words in that sequence. The target model must then generate a replacement token to keep the text flowing without stalling the generation pipeline.

However, the target model cannot simply sample from its original distribution. If it did, it would accidentally over-represent words that the draft model already successfully pushed through in previous parallel attempts, breaking the lossless guarantee.[4]

To perfectly balance the mathematical scales, the target model samples from a modified distribution called the residual probability. This distribution is calculated by subtracting the draft model's probability from the target model's probability for every possible word in the vocabulary.[1]

Specifically, the residual probability for any word is proportional to the maximum of zero or $p(x)$ minus $q(x)$. The system takes these remaining probability masses, normalizes them so they sum to 100 percent, and draws the replacement word from this newly formed pool.[1][3]

By exclusively sampling from the residual distribution upon rejection, the algorithm mathematically guarantees that the combined output of accepted draft tokens and residual replacement tokens perfectly matches the target model's exact distribution.[3]

When Draft Models Hallucinate

The elegance of the residual probability calculation becomes clear when examining extreme disagreements between the two models. Consider a scenario where the draft model is highly confident in a hallucinated word, assigning it a 90 percent probability.[3][4]

When a token is rejected, the target model samples exclusively from the residual probability mass to guarantee exact distribution matching.

If the target model evaluates that same word and assigns it only a 10 percent probability, the acceptance ratio drops to 11.1 percent. The token is almost certainly rejected during the verification phase, triggering the mathematical recovery mechanism.[1][4]

When the rejection occurs, the residual distribution must compensate. For the hallucinated word, subtracting 90 percent from 10 percent results in a negative number, which the formula clamps to zero. The rejected word is permanently eliminated from the replacement sampling pool.[4]

Meanwhile, the target model's preferred words, which the draft model ignored, will have large positive differences. The residual probability mass concentrates entirely on these neglected words, forcing the system to output the target model's true preference and correcting the hallucination.[4]

This mathematical correction is flawless, but it comes at a steep computational cost. When a token is rejected, the parallel speedup collapses for that cycle, reducing the system's output to just one token per forward pass and wasting the draft model's effort.

The Acceptance Rate Bottleneck

Because rejections destroy the speedup, the efficiency of speculative decoding is entirely governed by the acceptance rate. This metric tracks the percentage of drafted tokens that survive the target model's verification and make it into the final output.

"A well-matched draft model on the right type of content can achieve acceptance rates above 80 percent, leading to 2–3x speedups," explains a technical breakdown by MindStudio. "A poorly matched draft model might barely beat naive generation."

Illustration: Lossless acceleration allows AI providers to serve significantly more users on the same hardware footprint.

Surprisingly, making the draft model "smarter" does not always improve the acceptance rate. If a draft model is trained on different data than the target model, it might propose highly accurate words that the target model simply does not expect, triggering a costly rejection.[2]

To maximize the acceptance rate, engineers focus on alignment rather than raw intelligence. Techniques like distillation are used to train the draft model specifically to mimic the target model's quirks and probability distributions, ensuring they think in tandem.[2]

When the models are perfectly aligned, the residual probability approaches zero for most tokens. The draft model's distribution closely matches the target's, driving the acceptance ratio toward 100 percent and unlocking the maximum hardware speedup available to the system.[4]

Evolving Beyond Single Tokens

While standard speculative decoding evaluates one linear chain of guesses, newer frameworks construct complex trees of drafted tokens. Instead of guessing a single path, the draft model proposes multiple branching possibilities for the target model to verify simultaneously.[3]

This tree-based approach increases the likelihood that at least one drafted sequence will survive verification. However, it complicates the residual probability math, as a rejection on one branch requires the system to dynamically shift probability mass across the remaining valid branches.[3]

Recent advancements, such as Traversal Verification, attempt to solve this by evaluating sequence-level probabilities rather than isolating individual tokens. By traversing the draft tree from the leaf nodes back to the root, the algorithm preserves valid subsequences that older methods would have prematurely discarded.[3]

"Our approach considers the acceptance of the entire token sequence from the current node to the root," the Traversal Verification authors explain. This sequence-level math maintains the strict lossless guarantee while extracting even more speed from the same hardware.[3]

The Economics of Lossless Inference

The ability to accelerate large language models without altering their output has profound economic implications. Serving massive models requires vast clusters of expensive hardware, and inference costs scale directly with generation time.

By implementing speculative decoding, AI providers can serve two to three times as many users on the exact same hardware footprint. This reduces the energy consumption and capital expenditure required to run frontier models at scale.

By implementing speculative decoding, AI providers can serve two to three times as many users on the exact same hardware footprint.

"Producing results faster with the same hardware also means that fewer machines are needed for serving the same amount of traffic, which translates yet again to a reduction in the energy costs," noted Google Research in a retrospective.

The technique is particularly valuable for tasks with highly predictable outputs, such as writing code, formatting data, or extracting structured information. In these domains, the draft model can easily guess the syntax, driving acceptance rates higher and latency lower.

Ultimately, speculative decoding proves that the bottleneck in modern AI is not always the math itself, but how efficiently we move data through the silicon. By trusting a smaller model to draft the future, and relying on residual probabilities to catch its mistakes, the industry has found a way to cheat time.[4]

How we did this

Method
Recomputation of the residual probability distribution and expected token acceptance rates across varying draft-target divergence scenarios.
What we found
By computing the residual probability mass max(0, p(t) - q(t)), we find that a draft model's overconfidence (90% vs 10%) forces an 88.9% rejection rate, and the residual distribution concentrates 100% of its recovery mass on the target's preferred tokens. This mathematically guarantees the target distribution is restored, but reduces the effective token generation rate for that cycle to 1 token per forward pass, yielding zero speedup.
What we worked from
  • Draft model probability of 0.9 for a hallucinated primary token: 0.9 (90%) — OpenReview
  • Target model probability of 0.1 for the same token: 0.1 (10%) — OpenReview
  • Acceptance ratio defined as p(x)/q(x): p(x)/q(x) — arXiv
Limits of this analysis
This analysis assumes a static divergence for a single token and does not account for dynamic draft tree structures or feature-level drafting methods like EAGLE, which mitigate these extreme probability mismatches.

Terms to know

Autoregressive Decoding
The standard, sequential method of generating text where a model produces one token at a time based on all previous tokens.
Target Model
The large, highly capable language model that serves as the final authority on which words are accepted into the output.
Draft Model
A smaller, faster language model used to predict several upcoming words simultaneously for the target model to verify.
Rejection Sampling
A mathematical technique used to accept or reject a drafted word based on the ratio of its probability in both models.
Residual Probability
The remaining probability mass calculated by subtracting the draft model's confidence from the target model's confidence, used to sample a replacement word upon rejection.

Questions readers ask

Does speculative decoding reduce the quality of the AI's answers?

No. The mathematical verification step ensures that the final output is identical to what the large model would have produced on its own.

What happens if the draft model guesses completely wrong?

The target model rejects the incorrect guess and samples a replacement word from the residual probability distribution, though this eliminates the speedup for that specific cycle.

Why not just use the smaller draft model by itself?

Using the draft model alone would degrade the reasoning and writing quality of the output. Speculative decoding delivers the speed of the small model with the intelligence of the large model.

Different angles

Algorithmic Efficiency Researchers

Focus on maximizing acceptance rates and speedups through better draft-target alignment.

Researchers in this camp prioritize distillation and feature-level alignment to ensure the draft model closely mimics the target model. They argue that the true bottleneck in speculative decoding is not the verification math, but the frequency of rejections, which destroy the parallel speedup.

Production Infrastructure Engineers

Value the economic and latency benefits of lossless acceleration for serving massive models at scale.

For infrastructure teams, speculative decoding is primarily an economic tool. By tripling the effective throughput of a GPU cluster without compromising output quality, they can serve significantly more users on the same hardware footprint, drastically reducing the energy and capital costs of frontier AI.

Model Architecture Theorists

Explore alternative drafting mechanisms to bypass token-level probability bottlenecks.

Theorists argue that standard token-level drafting is inherently limited by the randomness of language. They advocate for sequence-level traversal and feature-level drafting, which evaluate the broader context of a generated sequence to preserve valid branches that simple rejection sampling would prematurely discard.

Algorithmic Efficiency Researchers 40%Production Infrastructure Engineers 35%Model Architecture Theorists 25%
Algorithmic Efficiency Researchers
Focus on maximizing acceptance rates and speedups through better draft-target alignment.
Production Infrastructure Engineers
Value the economic and latency benefits of lossless acceleration for serving massive models at scale.
Model Architecture Theorists
Explore alternative drafting mechanisms to bypass token-level probability bottlenecks.

Perspectives this story doesn't cover

  • Hardware Manufacturers
  • Open-Source Model Developers

Sources

Source coverage

4 outlets

3 viewpoints surfaced

Algorithmic Efficiency Researchers 40%Production Infrastructure Engineers 35%Model Architecture Theorists 25%
  1. [1]arXivAlgorithmic Efficiency Researchers

    Fast Inference from Transformers via Speculative Decoding

    Read on arXiv →
  2. [2]GMI CloudProduction Infrastructure Engineers

    How the Draft-and-Verify Loop Works

    Read on GMI Cloud →
  3. [3]OpenReviewModel Architecture Theorists

    Traversal Verification: A Lossless Speculative Decoding Algorithm

    Read on OpenReview →
  4. [4]Factlen Editorial TeamModel Architecture Theorists

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team →

Comments

Stay informed

Every angle. Every day.

Get Artificial Intelligence stories with full source coverage and perspective breakdowns, free every day.