Models & Algorithms•SOTAAZ Lab••KR

How Much Context Fits on an 8 GB GPU? Qwen3.5-9B From 8K to 256K, Against Five Real Tasks

We measured GPU memory for Qwen3.5-9B Q4_K_M at 8K to 256K context and counted the tokens of common tasks. On an 8 GB card one conversation gets 32K with the default KV cache and 64K with an 8-bit cache. A 20-turn coding chat reached 36,736 tokens, a 76-page paper 77,800, and coding agents send 648 to 15,970 tokens before you type a second message. Qwen3-8B, which keeps a cache in every layer, already needs 9.5 GiB at 32K.

How Much Context Fits on an 8 GB GPU? Qwen3.5-9B From 8K to 256K, Against Five Real Tasks

How Much Context Fits on an 8 GB GPU? Qwen3.5-9B From 8K to 256K, Against Five Real Tasks

Every local model guide gives you a memory number for one context length. In our guide to running Qwen locally, Qwen3.5-9B Q4_K_M used 5.8 GiB at 8K and 6.5 GiB at 32K, which put it on 8 GB cards. A reader replied on YouTube that 8K "should last about 30 seconds." Fair point. What a reader actually needs is two numbers side by side: how many tokens their task takes, and how much memory that much context costs.

So we measured both. GPU memory for the same 9B file from 8K to 256K, and the token counts of five things people do with a local model: a long chat, a PDF, a coding agent, a codebase, and the same text in Korean and English. We fixed the measurement plan, the expected numbers and the card thresholds before running anything.

How we measured

  • Runtime: llama.cpp 4da6337 (the same build as the earlier guide), llama-server -ngl 99 -fa on --parallel 1 -c <length>. One slot, so the whole context belongs to one conversation.
  • Models: Qwen3.5-9B Q4_K_M (Unsloth, the guide's file). For comparison, Qwen3-8B Q4_K_M, a similar-size model with full attention in every layer.
  • KV cache: the default f16, and 8-bit (-ctk q8_0 -ctv q8_0).
  • Memory: the server process's GPU memory, read every 0.2 seconds while a prompt filled 90% of the context. Each condition twice; the two readings were identical every time.
  • Tokens: counted with the 9B model's own tokenizer through the server.
  • Card limits: a card sold as 8 GB has 7.45 GiB. We leave 0.5 GiB for the desktop and driver, so "fits" means 6.95 GiB or less, and up to 7.45 GiB is "tight." The same rule gives 10.68 GiB for 12 GB cards and 14.40 GiB for 16 GB cards.

All memory numbers come from an A100. We did not run these on an 8 GB card; the memory a model takes is the same on any NVIDIA GPU with the same build and settings, but your desktop's share is your own.

Memory by context length

Line chart of peak GPU memory against context length on a log scale. Qwen3.5-9B Q4_K_M with f16 KV: 5.6 GiB at 8K, 6.4 at 32K, 7.4 at 64K, 9.5 at 128K, 13.6 at 256K. With q8_0 KV: 5.5, 6.0, 6.7, 8.1, 10.8. Qwen3-8B Q4_K_M with f16 KV: 6.1 at 8K, 9.5 at 32K; with q8_0: 5.5 and 7.4. Dashed lines mark the usable limit of 8 GB (6.95 GiB), 12 GB (10.68 GiB) and 16 GB (14.40 GiB) cards.
Context9B, f16 KV9B, q8_0 KVQwen3-8B, f16 KVQwen3-8B, q8_0 KV
8K5.62 GiB5.516.055.53
32K6.406.029.457.41
64K7.436.71beyond the modelbeyond the model
128K9.498.08beyond the modelbeyond the model
256K13.6210.83beyond the modelbeyond the model

On an 8 GB card, one conversation with the 9B gets 32K with the default cache, and 64K with an 8-bit cache. 64K at f16 is 7.43 GiB, just under the card's 7.45 GiB with nothing left for the desktop. A 12 GB card holds 128K either way, and a 16 GB card holds the model's full 256K.

Our 8K and 32K readings are a little lower than the guide's 5.8 and 6.5 GiB. The guide ran the server's default four slots; this time there is one.

Why the 9B grows slowly

The memory that grows with context is the KV cache: for every token, each attention layer keeps a key and a value. Qwen3.5-9B has full attention in only 8 of its 32 layers; the other 24 use a linear-attention layer whose state stays the same size however long the input gets. From the model's own metadata, 8 layers × 4 KV heads × (256 + 256) values × 2 bytes is 32 KiB per token. Qwen3-8B keeps a cache in all 36 layers: 36 × 8 × (128 + 128) × 2 bytes, 144 KiB per token, four and a half times as much.

We wrote those two numbers into the plan as predictions before measuring, and the f16 growth from 8K matched them within 0.25 GiB at every length (the plan allowed 0.3). From 8K to 32K the 9B added 0.77 GiB and Qwen3-8B 3.40 GiB. That is why, with an 8-bit cache, the same 8 GB card fits 64K for one model and only 8K for the other (Qwen3-8B at 32K is tight even then).

One more thing about Qwen3-8B: its training context is 40,960 tokens. When we asked for 64K, llama-server did not stop with an error. It printed a warning, quietly set the context to 40,960, and then rejected our 59K-token prompt. If you set -c above a model's training length, check the startup log for the context you actually got.

The 8-bit cache saves less than half

An 8-bit KV cache stores each value in a little over one byte instead of two, and the cache itself shrank exactly as expected: 4,096 MiB to 2,176 MiB at 128K. But total memory fell by only 1.4 GiB, not 1.9. The rest went to llama.cpp's compute buffer, which with an 8-bit cache grows with context length: 96 MiB at 8K, 696 MiB at 128K and 1,336 MiB at 256K, against 96, 216 and 344 MiB with f16. We found this in the server's detailed log after the 8-bit totals came out above our prediction, so it is a post hoc explanation, but the buffer sizes add up to the measured totals. In our KV cache quantization benchmark on Qwen3-8B, q8_0 showed no measurable quality loss, and decoding at 64K ran at 55% of f16 speed. So the 8-bit cache is the step that turns 32K into 64K on an 8 GB card. Its speed cost was measured on Qwen3-8B (55% of f16 at 64K); we did not measure speed on the 9B, which keeps a cache in only 8 of its 32 layers, so its cost may differ.

How many tokens do real tasks take?

TaskTokensSmallest context that holds it*8 GB card, 9B f168 GB card, 9B q8_08 GB card, Qwen3-8B f168 GB card, Qwen3-8B q8_0
Coding chat, after 10 turns15,49932Kfitsfitsnotight
Coding chat, after 20 turns36,73664Ktightfitsbeyond the modelbeyond the model
Same chat with reasoning on, after 10 turns34,67564Ktightfitsbeyond the modelbeyond the model
"Attention Is All You Need" PDF (15 pages)10,39232Kfitsfitsnotight
LoRA paper PDF (20 pages)19,84532Kfitsfitsnotight
Llama 2 paper PDF (76 pages)77,800128Knonobeyond the modelbeyond the model
nanoGPT model.py (331 lines)4,4038Kfitsfitsfitsfits
nanoGPT, all 15 Python files14,42632Kfitsfitsnotight
requests library, all 19 files in src/requests53,40764Ktightfitsbeyond the modelbeyond the model

\* Tokens plus 2,048 for the answer, rounded up to the next measured length.

The chat. We fixed 20 questions for a small expense-tracking app in Python and asked them in order, temperature 0, reasoning off. Answers ran 720 to 4,096 tokens and included code, so the conversation grew fast: 5,733 tokens after 5 turns, 15,499 after 10, 28,254 after 15 and 36,736 after 20. An 8K context fills in about seven turns of a chat like this; that YouTube comment was not far off. With reasoning on, we passed only the answers back, as chat apps do, so the reasoning itself never piled up; the answers came out longer instead: by turn 10 the conversation was 34,675 tokens, 2.2 times the reasoning-off run. Six of those ten answers hit our 4,096-token answer limit, so we stopped that run at turn 10.

PDFs. The 15- and 20-page papers we measured came to about 10K and 20K tokens. The 76-page Llama 2 paper is 77,800, and that is an undercount: the PDF reader skipped some of the text inside figures.

Code. One 331-line file is 4,403 tokens. A small project of 15 files fits in 32K, and a mid-size library like requests needs 64K.

Coding agents start before you do

Coding agents send a system prompt and tool descriptions with every request. We pointed four of them at a local server, gave each the same one-line task in an empty folder, and counted the tokens of the first request that carried the agent's tools:

AgentTokens in the first requestTools described
aider 0.86.2648none (edits are written as text)
Cline CLI 3.0.703,1214
OpenCode 1.18.357,40010
Qwen Code 0.25.015,97014

These are the starting cost of a session, not its size. Every file the agent reads and every tool result it gets back is added on top. Two agents also made side requests: OpenCode asked the model for a conversation title before starting, and Qwen Code ran a memory-saving step at the end. In Qwen Code's first request the tool descriptions alone were 35,697 characters, more than its 24,120-character system prompt. With Qwen Code, a 32K context is already half used by the first message. The numbers depend on each agent's version and settings, and they are not a ranking; aider spends fewer tokens partly because it gives the model no tools to call.

Korean takes about a fifth more

Every post on this blog has a Korean and an English version. Across 185 published pairs, the Korean version took a median 1.19 times the tokens of the English one (10th to 90th percentile, 1.10 to 1.28). Our Korean versions are written separately rather than translated, so part of this is length, not just the tokenizer. Still, as a rule of thumb, plan about 20% more context for the same material in Korean with this model.

What to set

  • On an 8 GB card with Qwen3.5-9B: -c 32768 for everyday use. Add -ctk q8_0 -ctv q8_0 and you can go to -c 65536 for long chats, coding agents or a mid-size codebase.
  • Run one slot (--parallel 1) if you are the only user, so the whole context belongs to your conversation.
  • A PDF as long as the 76-page Llama 2 paper needs 128K, which means a 12 GB card or splitting the document.
  • A model with full attention in every layer, like Qwen3-8B, needs over four times the memory per token. On 8 GB it holds 8K comfortably and 32K only with an 8-bit cache, at 7.41 GiB, which is tight.

Limits

  • Memory was measured on an A100, not on an 8 GB card. Leave room for your own desktop.
  • We measured whether a context fits, not how well the model answers at that length.
  • One chat, written by us, with one model at temperature 0. Other conversations grow at different rates.
  • The agent numbers are first requests only, for one task, with each agent's default settings.
  • Qwen3-8B stops at 32K in our table because of its training context, so we have no comparison above that.
  • The compute-buffer explanation for the 8-bit cache comes from logs we read after the results.

The plan, scripts, per-run logs and results are in the reproduction package below. For what an 8-bit cache does to quality and speed, see our KV cache quantization benchmark.

Files for this post

Reproduction package

The harness, raw logs and result CSVs behind this post. No account needed; the README has a command that recomputes the tables without a GPU.

ctx-budget-8gb-repro.zip · 279 KB

Download