llama.cpp KV Cache Quantization: Which -ctk and -ctv to Use, Measured on the Current Build
Keep f16 if the KV cache fits. If not, set -ctk q8_0 -ctv q8_0: it saved 4.2 GiB at 64K but decoded at 56% of f16's speed there.

llama.cpp KV Cache Quantization: Which -ctk and -ctv to Use, Measured on the Current Build
Short answer. If the KV cache fits in VRAM, leave it at the default f16. If it does not, use -ctk q8_0 -ctv q8_0. Set both flags to the same value, and do not expect the memory to be free: at 64K tokens of context the quantized cache decodes at roughly half the speed of f16.
This page is the short version of three measurement posts on this blog, rechecked on the llama.cpp build of 4 October 2026. The recheck mattered. A change merged on 9 September altered what happens with the less common types, and part of what I published in September no longer holds on current builds. Where a number below comes from an older build, it says so.
The commands
# server, chat, CLI: the same two flags
llama-server -m model.gguf -ngl 99 -fa on -ctk q8_0 -ctv q8_0 -c 32768
# compare types on your own card before you commit to one, one type per run
for t in f16 q8_0 q4_0; do
llama-bench -m model.gguf -ngl 99 -fa on -ctk $t -ctv $t -p 4096 -n 128 -d 0 -d 32768
done-ctk and -ctv are the short forms of --cache-type-k and --cache-type-v. Keep -fa on: quantized V caches need flash attention in llama.cpp. llama-bench also accepts comma-separated lists, but -ctk f16,q8_0,q4_0 -ctv f16,q8_0,q4_0 runs all nine K/V combinations, six of them the mixed pairs this page advises against. The loop keeps K and V equal.
What each setting costs
Qwen3-8B Q4_K_M on one A100 80GB. Speeds are tokens per second on the 4 October build (b11390 (dd26678), default CUDA options). Memory is the cache alone at 32K tokens, computed from the type's bits. Perplexity is from the September build (b10753 (69320fe)); the cache format is set by the type, so it should not move with the kernels, but I did not re-run it.
-ctk / -ctv | Cache at 32K | Perplexity vs f16 (4K–32K) | Prefill 4096 | Decode at 64K depth |
|---|---|---|---|---|
| f16 / f16 | 4.50 GiB | — | 4,704 | 79.6 (100%) |
| q8_0 / q8_0 | 2.39 GiB | −0.0% to +0.1% | 4,584 | 44.3 (56%) |
| q4_0 / q4_0 | 1.27 GiB | +0.5% to +1.1% | 4,599 | 42.4 (53%) |
| bf16 / bf16 | 4.50 GiB | not measured | 4,672 | 41.3 (52%) |
| q5_1 / q5_1 | 1.69 GiB | not measured | 4,583 | 27.8 (35%) |
| q8_0 / q4_0 | 1.83 GiB | not measured | 4,582 | 33.8 (42%) |
| iq4_nl / iq4_nl | 1.27 GiB | not measured | 104 | not measured |
Three things in that table decide most setups.
Prefill barely moves; decode at depth does. Every working type prefills within 3% of f16. The cost shows up when the cache is already long, because each new token reads all of it. With 8K tokens in the cache, q8_0 decoded at 83% of f16; at 64K, 56%. The memory you saved is paid back in time at exactly the context length where you needed the memory.
q8_0 and q4_0 decode at nearly the same speed, within 5% of each other at 8K, 32K and 64K. Between them the choice is memory against quality: q4_0 holds about 1.1 GiB less at 32K for this model and costs about 1% perplexity. Take q4_0 only when that gigabyte decides whether the model loads.
The memory saving is real. Peak VRAM during the 64K runs fell by 4.19 GiB with q8_0 and 6.45 GiB with q4_0. The arithmetic predicts 4.22 and 6.47, so the measured peak matched it to within 0.03 GiB.
Types that work but you should not pick
On the September build, four of these types and every mixed K/V pair prefilled at 83 to 284 tokens per second against about 4,600, because the default CUDA build compiled no flash-attention kernel for them. I published that as "set both flags to the same value, and use only four types."
On the current build that cliff is gone for most of them. A change merged on 9 September (llama.cpp #28079) did two things. It made q4_1, q5_0 and q5_1 supported types for flash attention. And when the decode kernel for a K/V pair was not compiled, it now converts K and V to f16 on the fly instead of failing over. llama-server prints one line when that happens:
ggml_cuda_flash_attn_ext_vec: no FlashAttention vector kernel compiled for K/V types q5_1-q5_1, converting K and V to f16 instead (slow). Add "q5_1-q5_1" to GGML_CUDA_FA_QUANTS to compile it.llama-bench does not show that line unless you pass -v, because it silences the library's log by default.
So q5_1 and the q8_0 / q4_0 mix now prefill at full speed, and each gave a normal answer to one short test request on llama-server; I did not measure their quality. They are still a bad trade on a default build. q5_1 decodes at 35% of f16 at 64K, slower than q4_0, while holding more memory than q4_0. The mix sits between them on speed and saves less than q4_0. bf16 is the other trap: it holds exactly as much as f16 and decodes at 52% of it at 64K, so it buys nothing on this card.
iq4_nl is the one type that is still unusable. It is not on the flash-attention list at all, and it prefilled at 104 tokens per second.
If you have a reason to want q5_1 or a mixed pair, compile its decode kernel in: -DGGML_CUDA_FA_QUANTS="q8_0-q8_0;q4_0-q4_0;f16-f16;bf16-bf16;q5_1-q5_1", or all at the cost of a much longer build. The old option GGML_CUDA_FA_ALL_QUANTS is now a deprecated alias for all. I have not measured a build with the extra kernels on the current commit.
On a server, output length decides the cost
Single-stream decode is the worst case. In the llama-server measurement on the September build, with a 32K prompt, q8_0 cost 9% of total throughput when each request produced 128 tokens and 22% when it produced 1,024. Prefill takes most of a short-answer request and quantization barely touches prefill. If your answers are long relative to your prompts, measure at your own ratio.
Before you trust a number from your own card
- Check that nothing else is on the GPU.
nvidia-smi --query-compute-apps=pid,used_memory --format=csv. My first sweep in September readf16prefill at 1,976 tokens per second instead of 4,707, because a run I thought I had stopped was still sharing the card. - Measure at the depth you will run. At an empty cache every working type looks the same. Add
-d 32768(or your real context) to llama-bench. - Watch the log, not just the speed. The
converting K and V to f16line is the only sign that a pair fell off its compiled kernel.
How big the cache is for your model
Per token, the cache holds layers × KV heads × head dimension × 2 (K and V) × bytes per value. Qwen3-8B has 36 layers and 8 KV heads of dimension 128, which is 144 KiB per token at f16, or 4.5 GiB at 32K. q8_0 stores 8.5 bits per value and q4_0 4.5 bits. A model with more KV heads has a proportionally larger cache and, plausibly, a larger slowdown at depth; I have measured only this one. The general formula and a script for other models are in KV Cache Explained.
What this does not show
One model, one weight quantization, one A100, batch size one for the speed table. Perplexity and the server numbers come from the September build, not the current one. The 64K decode figures are a single run each. llama.cpp changes weekly, and the 9 September change is a reminder that a table like this one can go stale in a month; the date and commit are in the note below so you can tell.
The three posts behind this page have the full tables: single-stream speed, memory and perplexity, llama-server under concurrent load, and every -ctk value on the September build. The sweep scripts are in the free KV cache benchmark harness.
Current-build measurements: llama.cpp b11390 (dd26678) (ggml-org master, 4 October 2026), CUDA build for sm80 with default options (GGML_CUDA_FA_QUANTS=q4_0-q4_0;q8_0-q8_0;f16-f16;bf16-bf16), Qwen3-8B Q4_K_M, one A100 80GB PCIe, -fa on -ngl 99. Prefill: -p 4096 -n 128 -d 0 -d 8192 -r 2. Depth: -p 0 -n 128 -d 32768 -d 65536 -r 1, VRAM sampled every 0.5 s with nvidia-smi. Warning lines from llama-server with -c 8192 and one short request. Scripts scripts/bench-kv-types-master.sh and scripts/bench-kv-depth-master.sh; raw JSON and logs in drafts/llamacpp-kv-guide/. Perplexity and server figures: llama.cpp b10753 (69320fe) (2 September 2026), same card and model, from the posts linked above.
New posts on this topic go into the weekly newsletter. One email a week at most.
Courses that go with this post
SOTAAZ courseCourses are sold one at a time. If you want all 13, there is a $199 lifetime bundle.
Subscribe to Newsletter
Related Posts

llama.cpp KV Cache Quantization, Measured on One A100 — q8_0 Is Free at 4K and Costs Half Your Decode Speed at 64K
Mainline llama.cpp, Qwen3-8B, one A100: -ctk q8_0 -ctv q8_0 matches f16 perplexity and cuts the 32K cache by 2.1 GiB, but decode at 64K depth drops to 55% of f16 (q4_0: 50%). Two other settings, q5_1 and a q8_0/q4_0 mix, silently ran prefill on the CPU at 43 and 63 tokens per second.

TurboQuant in vLLM on One A100 — Capacity, Speed, and Accuracy of All Four Presets on an 8B Model
vLLM 0.28, Qwen3-8B bf16, one A100 80GB: KV capacity, batched throughput, 32K decode, needle-in-haystack, and GSM8K for bf16, fp8, and all four TurboQuant presets — the 8B size vLLM's own study skipped.

TurboQuant From Scratch on Real KV Tensors — What 3 Bits Actually Cost, and Why the Forks Beat the Paper's Layout
PolarQuant in 60 lines of PyTorch on real KV from Llama-3.2-1B and Qwen3-8B: 3-bit costs +10% perplexity, k8v4 +0.2%, QJL only pays below 4 bits, and the block-32 layout explains half the forks' edge.

