Models & Algorithms•SOTAAZ Lab••KR

Speculative Decoding Speedups Depend on the Task: Prompt Lookup 6.2× on a File Edit, DFlash 0.7× on an Essay (Qwen3.5-9B, llama.cpp)

We measured three kinds of speculative decoding in llama.cpp on one Qwen3.5-9B file, alone on an A100: prompt n-gram lookup, the model's own MTP head, and a DFlash draft model. Rewriting a 130-line file ran 6.2× faster with n-gram lookup and 3.2× with DFlash. On a 500-word essay, MTP was 4% faster, n-gram lookup slightly slower and DFlash 0.69×. At temperature 0, n-gram and MTP output matched the plain run token for token; DFlash output diverged on three of four tasks.

Speculative Decoding Speedups Depend on the Task: Prompt Lookup 6.2× on a File Edit, DFlash 0.7× on an Essay (Qwen3.5-9B, llama.cpp)

Speculative Decoding Speedups Depend on the Task: Prompt Lookup 6.2× on a File Edit, DFlash 0.7× on an Essay (Qwen3.5-9B, llama.cpp)

Local inference servers advertise speedups like "up to 4× faster." The Rapid-MLX README is a good example of reading those numbers honestly: on Qwen3.5-9B 4-bit on an M4 Pro, its decode speed was 1.50× at the median of 18 tasks (lowest 1.28×), up to 4.27× on one task, and a long agent turn was close to a tie at 1.05×. The README also says where the gain comes from: speculative decoding.

Speculative decoding works like this. Something cheap guesses the next several tokens, and the real model checks all the guesses in one pass, keeping the ones it would have produced anyway. When the guesses are good, you get several tokens for the price of one model step. When they are bad, you pay for the guessing and get nothing.

So the speedup should depend on how guessable the output is, which depends on the task. We measured that in llama.cpp with three kinds of guesser on the same model file. We fixed the tasks, conditions and decision rule before measuring.

How we measured

  • Runtime: llama.cpp 4da6337, llama-server -ngl 99 -fa on --parallel 1 -c 16384 --jinja, on one A100 80GB with no other process on the GPU. We logged every process on the GPU every two seconds; runs that overlapped another job were measured again.
  • Model: Qwen3.5-9B converted from the original weights to Q4_K_M ourselves, so the file keeps the model's built-in MTP layer (the Unsloth GGUF in our earlier posts does not include it). Every condition uses this one file.
  • Four conditions:

- Off: no speculative decoding.

- Prompt n-gram (--spec-type ngram-mod): looks for the last few tokens earlier in the context and proposes whatever followed them there. No extra model.

- MTP (--spec-type draft-mtp): Qwen3.5's own multi-token prediction layer guesses ahead.

- DFlash (--spec-type draft-dflash): a separate 6-layer draft model, z-lab/Qwen3.5-9B-DFlash, that proposes a block of up to 15 tokens at once.

  • Four tasks: rewrite a 130-line Python file adding type hints and change nothing else; write a new to-do app; solve three GSM8K math problems step by step; write a 500-word essay.
  • Runs: three per task and condition, at temperature 0 and 1.0, reasoning off, up to 1,024 output tokens. We report decode speed from the server's timings, not prompt processing.
  • Rule: a condition counts as faster or slower than off only if all three of its runs are faster, or all slower, than all three off runs.

The result

Horizontal bar chart of decode speed relative to speculative decoding off, temperature 0. Rewrite a 130-line file: prompt n-gram 6.19×, MTP 1.78×, DFlash 3.19×. Write new code: 0.97×, 1.59×, 2.27×. Math word problems: 0.98×, 1.52×, 1.74×. Essay: 0.99×, 1.04×, 0.69×. Off ran at 126 tokens per second. Prompt n-gram values are from a fresh server per request, measured after the first results.
Task (temperature 0)Prompt n-gram (fresh server)MTPDFlash
Rewrite a 130-line file6.19× faster (781 tok/s)1.78× faster3.19× faster
Write new code0.97× slower1.59× faster2.27× faster
Math word problems0.98× slower1.52× faster1.74× faster
500-word essay0.99× slower1.04× faster0.69× slower

"Faster" and "slower" are the verdicts of our rule: all three runs on the same side of all three plain runs. Without speculative decoding the model generated about 126 tokens per second on every task. Temperature 1.0 gave the same pattern: 6.25×, 1.78× and 3.35× on the file rewrite, and 1.00×, 1.04× and 0.66× on the essay.

The task decided the speedup more than the method did. On the file rewrite, where the answer copies most of the input, every method helped and n-gram lookup helped most. On the essay, MTP was 4% faster (131 against 126 tokens per second), n-gram lookup was slightly slower, and DFlash was much slower. We had predicted this order before measuring, and it held for every method.

Why each method behaves the way it does

The server reports how many guessed tokens it drafted and how many it accepted, which explains the pattern.

  • Prompt n-gram can only propose text that already appears in the context. On the file rewrite it accepted every one of its 984 guesses per run, because the output is mostly the input. On new code and math it found almost nothing to copy (0 to 4 accepted tokens per run), so it cost a little and gained nothing: slightly slower in every run.
  • MTP guesses from the model's own extra layer and was accepted 98% of the time on the rewrite, 84% on new code, 78% on math and 41% on the essay. It drafts only a few tokens per step, so its gains stay modest, 1.5× to 1.8× on the first three tasks and 1.04× on the essay, and it never made anything slower.
  • DFlash drafts up to 15 tokens per step with a separate model. When its guesses land, that pays off more than MTP: 2.27× on new code. On the essay only 8% of its drafted tokens were accepted, and the cost of running the draft model every step outweighed the few tokens it saved: 0.69×.

So the best method depends on the job. For editing existing code or text, prompt n-gram is free and fastest. For writing new code, a trained draft like DFlash helped most. For prose, MTP gave a small gain and was the only method that never slowed anything down; DFlash is better left off.

Is the output the same?

Speculative decoding is supposed to produce the same text as the model alone, because the model checks every guess. We checked this at temperature 0 by comparing tokens. First, two plain runs against each other: identical on all four tasks, so the GPU itself was not adding noise. Then each method against the plain run:

TaskPrompt n-gramMTPDFlash
Rewrite a fileidenticalidenticalidentical
New codeidenticalidenticaldiffers from token 206 (first 23% the same)
Mathidenticalidenticaldiffers from token 83 (first 9% the same)
Essayidenticalidenticaldiffers from token 12 (first 2% the same)

N-gram and MTP reproduced the plain output exactly. DFlash did not on three of the four tasks, so its speedups there were measured on different text of different length: 829 tokens against 898 on new code, 627 against 894 on math, 681 against 611 on the essay. Decode speed is per token, so the comparison still holds, but the content changed. On the math task, where DFlash's answer was 30% shorter, it got all three problems right while the plain run and MTP got the third one wrong ($25,000 instead of $70,000); one prompt is no evidence either way about quality, only that the text was not the same. We did not trace the cause, so we are not claiming DFlash is wrong in principle. If identical output matters to you, check it on your own prompts.

Two things that changed our numbers

Another job on the GPU. Right after measurement started, a process from another session ran on the same GPU for a short while. The three plain runs of the file rewrite came out at 48 to 56 tokens per second; measured again on a quiet GPU they were 126. Our runner only checked the GPU before each request, so after that we logged the GPU every two seconds and remeasured all eight runs that had finished before the log started. If your own speedups look strange, check what else is on the card.

The n-gram pool survives between requests. Our first n-gram runs used one server for all requests, and repeated identical requests got faster each time: 123, then 195, then 307 tokens per second on the new-code task. ngram-mod keeps its table of seen n-grams for the life of the server, and at temperature 0 the second request's answer is the same as the first, so it copies its own earlier output. Averaged over the three runs, that gave:

Task (temperature 0)Prompt n-gram, one server for all runsPrompt n-gram, fresh server per run
Rewrite a file7.20× (faster)6.19× (faster)
New code1.65× (runs too uneven to call)0.97× (slower)
Math1.85× (runs too uneven to call)0.98× (slower)
Essay1.02× (runs too uneven to call)0.99× (slower)

Real use, with a new prompt each time, looks like the right-hand column, so that is what the chart shows. We added the fresh-server runs after seeing the left-hand column. The left-hand column is useful too: if you really do send the same kind of request over and over, such as reformatting similar files, n-gram lookup can get faster as the server warms up.

What to set

  • Editing code or text, or anything where the answer repeats the input: --spec-type ngram-mod. No extra model, and 6× on our file rewrite.
  • Writing new code: a DFlash draft trained for your model, if one exists (-md <draft.gguf> --spec-type draft-dflash). Check the output against a plain run if you need identical results.
  • Prose or chat: MTP if your GGUF includes the MTP layer; expect a few percent at most. Leave DFlash off.
  • Before trusting an advertised speedup, ask what task it was measured on.

Limits

  • One model (Qwen3.5-9B Q4_K_M), one runtime build, one GPU. Speedups on a Mac or a consumer GPU will differ in size; we expect the order across tasks to stay the same.
  • Four tasks, one prompt each, three runs per condition. The math and essay tasks are short, single-turn prompts.
  • Each method ran at its default or documented settings; tuning draft lengths could change the numbers.
  • Our quantized file was made without an importance matrix, unlike the Unsloth file in our earlier posts, so speeds are not directly comparable with those.
  • The fresh-server n-gram runs were added after we saw the repeated-request effect.
  • We did not investigate why DFlash output diverged.

The plan, scripts, per-run results and server logs are in the reproduction package below.

Files for this post

Reproduction package

The harness, raw logs and result CSVs behind this post. No account needed; the README has a command that recomputes the tables without a GPU.

spec-decode-by-task-repro.zip · 458 KB

Download