With one user, a GPU spends most of each decode step waiting for the weights to arrive from memory, and its math units sit mostly idle. Batching fills that idle math with other users’ tokens. Speculative decoding fills it with guesses: something cheap guesses the next few tokens, and the model checks all of them in one step. When the guesses are right, one step produces several tokens. This post measures how often the guesses are right, and what that does to speed on an H100.

How it works Link to heading

A decode step makes one token, and it reads all the weights to do it. Checking five tokens in a step reads the same weights. The model only has to do the math for five positions instead of one, and with one user there’s plenty of spare math. So the step costs about the same whether it checks one token or five.

Speculative decoding uses that. Each step:

  1. A guesser proposes the next few tokens, here 4.
  2. The model runs one step over all of them, and at every position works out what it would have picked itself.
  3. It keeps the guesses up to the first one it disagrees with, and adds its own pick at that point.

So every step produces at least one token, the model’s own, plus every correct guess before it. With greedy decoding, the kept tokens are the ones the model would have picked anyway, so the answer doesn’t change, only how many steps it takes.

The guesser can be anything cheap. I tried two that vLLM supports:

  • A draft model: a much smaller model from the same family, here Qwen3-0.6B guessing for Qwen3-8B. It runs 4 times per step, once per guessed token.
  • N-gram lookup: no model at all. It looks for the last few tokens earlier in the prompt and guesses whatever followed them there. It only works when the answer repeats the prompt.

Setup Link to heading

  • Qwen3-8B in BF16 on one H100 80GB on Modal, vLLM 0.30, greedy decoding, thinking off, up to 256 output tokens.
  • Two tasks with real text. Math: GSM8K test questions, where the answer is new text. Repeat: a WikiText passage the model is asked to repeat and then extend by one sentence, where the answer copies the prompt.
  • 1, 8 and 64 users at once, with 16, 48 and 192 requests.
  • Acceptance comes from vLLM’s own counters for guessed and accepted tokens.
  • The script and raw results are in the lab repo.

The whole run happened three times, each in a fresh container. For the setups in the tables below, runs 2 and 3 agreed within 7%, and the tables are run 3. Run 1 was much slower for every speculative setup; the caveats explain it.

How often the guesses are right Link to heading

                  share of guesses accepted     tokens accepted per step
                  math       repeat             math       repeat
n-gram lookup     28%        86 to 90%          1.1        3.4 to 3.6
draft model       72 to 75%  80 to 84%          2.9 to 3.0 3.2 to 3.4

The draft model guesses well on both tasks, because it’s a smaller version of the same model and tends to pick the same words. N-gram lookup is right 9 times in 10 when the answer copies the prompt, and wrong most of the time on math, where the answer is new.

How much faster Link to heading

Tokens per second for one user, out of the 4 guesses a step:

                       math                          repeat
users            none   draft   n-gram          none   draft   n-gram
    1             154     255      159           153     265      373
    8             149     214      147           145     249      337
   64             124     157      109           111     140      172

With one user, the draft model made math 1.66 times faster (255 / 154), and n-gram lookup made repeating 2.44 times faster (373 / 153). N-gram lookup did nothing for math with one user, and with 64 users it made math slower than no speculation at all.

The gains are smaller than the acceptance numbers suggest, because a speculative step isn’t free. With one user, a normal step takes 1000 / 153.8 = 6.5 ms. For the draft model on math:

tokens per step     2.88 accepted + 1 of the model's own  =  3.88
step time           3.88 tokens / 255.4 tokens per second =  15.2 ms

Each step makes almost 4 tokens but takes 2.3 times as long as a normal step: the draft model runs 4 times, and the model checks 5 positions. N-gram lookup on the repeat task makes 4.59 tokens a step in 12.3 ms. Speculation wins when enough guesses land to pay for a step that costs about twice as much.

Why it fades with more users Link to heading

Speculation spends spare math, and batching spends the same spare math. With more users in each step, there’s less of it left for checking guesses. The speedup for one user’s tokens shrinks as users are added:

draft model on math     1 user: 1.66x     8 users: 1.44x     64 users: 1.27x
n-gram on repeat        1 user: 2.44x     8 users: 2.32x     64 users: 1.56x

That’s why speculative decoding matters most where a few users want fast answers, and least where a provider packs a GPU with users to bring down the cost per token.

Caveats Link to heading

  • Run 1 was much slower. In the first run, every speculative setup was 1.9 to 2.5 times slower than in runs 2 and 3, and plain decoding was 4 to 11% slower, with the same acceptance. vLLM logs “Async scheduling not supported with ngram-based speculative decoding and will be disabled”, so with speculation, the CPU’s work for each step isn’t overlapped with the GPU’s. That makes speculation sensitive to the host CPU, and that container’s CPU was slower. Measure on your own hardware.
  • Same answers, mostly. With greedy decoding, speculation should give exactly the text the plain server gives. In practice some answers differed, for example 15 of 16 math answers matched for the draft model with one user, because small differences in floating-point arithmetic change an occasional token.
  • One model, two tasks, 4 guesses a step. Other draft sizes, other guess counts and methods that train a draft head for one model (EAGLE) behave differently. vLLM’s GPU version of n-gram lookup made similar speeds with fewer tokens accepted per step.