Strata vs llama.cpp on the Same Qwen3.8-Flash-Next File: No Score Difference Found, Faster Within 12 GiB
Same IQ2_XS file in both engines: GSM8K 93.3% vs 93.0%, HumanEval 157/164 each, no vision gap showed up on 100 COCO images. Held to 12 GiB, Strata (draft decoding on) decoded 62.9 tok/s to llama.cpp's 23.4.

Strata vs llama.cpp on the Same Qwen3.8-Flash-Next File: No Score Difference Found, Faster Within 12 GiB
Strata reached the front page of Hacker News on 4 October under the title "Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s". It is an inference engine built for one model, Qwen3.8-Flash-Next, and it promises to run that model on a gaming card with 12 GB of memory. The thread split into two camps. Some people posted speed numbers well above what they got from llama.cpp. Others said the answers felt worse, and one commenter reported that on 50 images, with the same GGUF and vision adapter, Strata's median miss when pointing at objects was more than three times llama.cpp's.
Both claims can be tested, because the model file is a standard GGUF that also runs in llama.cpp. So I ran the same file in both engines, with the same requests, and compared scores, repeat consistency and speed.
The short answer
- No score difference showed up. On the same IQ2_XS file, GSM8K was 280/300 in Strata and 279/300 in llama.cpp, and HumanEval was 157/164 in both. Pointing at objects in 100 COCO images, the median miss was 37.6 pixels in Strata and 37.7 in llama.cpp. The three-times vision gap from the thread did not show up on these images.
- Strata was faster, by more when the GPU was small. With the card held to about 12 GiB, Strata decoded 62.9 tokens per second and llama.cpp 23.4. A 29K-token prompt was read at 2,726 and 275 tokens per second.
- One real difference. With the GPU held to 12 GiB, Strata's default settings gave a byte-identical answer on all three repeats for only 12 of 30 questions at temperature 0. llama.cpp gave identical answers on all 30. Strata's documentation describes this and offers switches that fix it. They cost about a quarter of the decode speed.
My hardware is an A100 80GB and a 64-core server CPU, not a gaming PC. The scores and the repeat result carry over. The speed numbers do not, as explained below.
Why a 125B model fits next to a 12 GB card
The number 125B describes storage, not work. According to the Qwen model card, Qwen3.8-Flash-Next has 125B parameters of which 6B are active for each token. On top of that come a 51B-parameter n-gram embedding table and a 4B-parameter MTP layer for drafting tokens ahead. Each layer has 512 experts, and each token uses 10 of them plus one shared expert.
That split decides where each piece can live. The n-gram table is a lookup: each token reads a few rows, so it can stay in RAM or on an SSD. The experts are most of the remaining weight, but any one token touches only a few, so they can sit in system RAM. Only the parts every token uses need the GPU. The quantized files I used (ISTA-DASLab's GSQ-RCO release) ship in two shards that follow this split: 39.2 GB of weights and a 28.8 GB n-gram table.
The two engines handle the experts differently. In llama.cpp, the usual way to fit a small card is --n-cpu-moe N, which keeps all experts of the first N layers on the CPU. Strata ranks individual experts by how often they are used, keeps the most-used ones in VRAM, and computes the rest on the CPU or streams them over PCIe. In my 12 GiB score run its log reported, per request, a median 77% of expert reads served from the GPU cache (44% to 95% across 490 requests) and another 17% streamed over PCIe. That difference in placement, together with drafting several tokens at once, is where the speed below comes from. The model's structure is what makes either approach possible.
What I ran
Both engines got the same model file, checked against the SHA-256 hashes on Hugging Face, and the same request bytes over their OpenAI-compatible APIs: temperature 0, thinking off, one request at a time.
| Strata | llama.cpp | |
|---|---|---|
| Version | commit 6f32ec0 (engine 0.1.39), Docker image built for sm80 | master d89651a (5 Oct 2026), CUDA 12.1 |
| Settings | what its setup wrote: KV cache int8, MTP drafting on (--spec 4) | -fa on -c 65536, KV cache f16, no drafting (this GGUF has no MTP layer) |
| Condition A | every expert on the GPU (setup default on an 80 GB card) | -ngl 99 |
| Condition B | --vram-reserve-mib 69000: peak 12.3 GiB used (12.25 in the speed run) | -ngl 99 --n-cpu-moe 42: peak 12.3 GiB used (12.19 in the speed run) |
I wrote the questions, the sample sizes and the decision rules down and committed them before the first run. Condition B for the score tests was not in that plan. I added it after condition A, because in condition A every expert sat on the GPU, so the CPU path that a 12 GB card actually uses had not been tested. The server's disk is a spinning drive, so Strata's setup kept the n-gram table in RAM. I gave llama.cpp the same (--lazy-mode off) in the condition B score run and in every speed run. Only the condition A score run used the model card's --lazy-mode on, which affects speed, not answers.
Scores: no difference I could detect
| Condition | llama.cpp | Strata | Only llama.cpp right | Only Strata right | McNemar p | |
|---|---|---|---|---|---|---|
| GSM8K (300) | A | 279 (93.0%) | 280 (93.3%) | 2 | 3 | 1.00 |
| GSM8K (300) | B | 280 (93.3%) | 281 (93.7%) | 3 | 4 | 1.00 |
| HumanEval (164) | A | 157 (95.7%) | 157 (95.7%) | 1 | 1 | 1.00 |
The two engines did not write the same text. In condition A only 138 of 300 GSM8K answers matched word for word, but 290 ended on the same number. llama.cpp's own answers changed about as much between its two conditions: 156 of 300 identical, 291 with the same number. Different kernels round differently, and over a few hundred tokens the wording drifts. In this sample, the final answers rarely did.
A caution on what these numbers can show. With 300 questions, a gap of one or two points would not come out as significant. GSM8K and HumanEval are also old and easy enough that this model scores above 90% on them. A difference in harder, longer work, especially with thinking on, is outside this test.
The vision gap did not reproduce on public images
The commenter asked for the coordinates of a named object in 50 images and measured the distance to the right answer. Strata missed by a median 154.8 pixels and llama.cpp by 46.5. Their images are not public, so I built a public version: 100 COCO val2017 images, each with one object of a category that appears only once in it, and the question "Point to the {object} in the image" with the answer in the model's usual 0-1000 coordinates. In a later reply the commenter said they used the BF16 vision file and the same 0-1000 format, so those two parts match my setup.
| Condition | Engine | Inside the box (of 100) | Median miss | Mean miss |
|---|---|---|---|---|
| A | llama.cpp | 95 | 37.7 px | 46.3 px |
| A | Strata | 93 | 37.6 px | 48.0 px |
| B | llama.cpp | 95 | 36.7 px | 46.1 px |
| B | Strata | 95 | 37.7 px | 46.0 px |
For every image, both engines produced the same number of prompt tokens (129 to 439 per image, median 298). Of Strata's two fewer hits in condition A, one was an answer with no coordinates ("There is no bowl in the image."). In condition B the hit counts were equal. The Strata documentation says the same thing: "Answers match llama.cpp's multimodal implementation token for token on our test images."
That does not prove the commenter wrong. Their images, prompt and Strata version all differ from mine, and Strata's changelog shows its image handling changing more than once in recent releases (0.1.23, 0.1.33). What I can say is that on this public set, with this version, the gap does not appear.
Same question, different answer
I asked the first 30 GSM8K questions three times each, at temperature 0, in each setup.
| Setup | Questions with three identical answers | Same final number |
|---|---|---|
| llama.cpp, A and B | 30 / 30 | 30 / 30 |
| Strata, A (all experts on the GPU) | 30 / 30 | 30 / 30 |
| Strata, B, defaults | 12 / 30 | 29 / 30 |
| Strata, B, reproducibility switches | 30 / 30 | 30 / 30 |
Strata's documentation explains why. When experts are split between GPU and CPU, the same expert can be computed by different kernels, which round differently. Which kernel runs depends on the drafting and on which experts moved into VRAM since the last request. The fix it documents is STRATA_IQ_MT_MIN=1 plus --prompt-cache 0 --adapt-swaps 0 --pcie-frac 0. With those set, all 30 repeated exactly. Decode speed in condition B went from 62.9 to 46.3 tokens per second.
For chat this hardly matters: one of 30 final answers changed. For A/B tests, cached evaluations or anything that compares two runs, it does. Two runs of the same prompt can differ for reasons unrelated to what you changed.
Speed
Fifteen different 256-token answers for decode, and three prompts each of about 4K and 29K tokens for reading, every request different so no prompt cache could help. Speeds are medians timed by the client.
| Condition | Engine | Decode tok/s | Read 4K prompt tok/s | Read 29K prompt tok/s |
|---|---|---|---|---|
| A, whole card | llama.cpp | 63.0 | 879 | 965 |
| A, whole card | Strata | 110.7 | 1,995 | 2,826 |
| B, 12 GiB, 64 cores | llama.cpp | 23.4 | 242 | 275 |
| B, 12 GiB, 64 cores | Strata | 62.9 | 1,865 | 2,726 |
| B, 12 GiB, 8 cores | llama.cpp | 8.8 | 237 | 274 |
| B, 12 GiB, 8 cores | Strata | 43.1 | 1,763 | 2,736 |
Within 12 GiB, Strata decoded 2.7 times as fast as llama.cpp on all cores and 4.9 times on eight. It also read long prompts at almost the same speed with 12 GiB as with the whole card, while llama.cpp dropped to under a third.
Three things in this table favour Strata, and only the last is a choice I made. Its drafting layer is on and llama.cpp has none for this file. Its KV cache is 8-bit, against llama.cpp's 16-bit. And --n-cpu-moe is the simple way to fit a small card, not the only one. I did not try llama.cpp's per-tensor overrides (-ot), and I did not run Strata with drafting off.
The hardware also matters. An A100 has far more memory bandwidth than a gaming card, and the EPYC 7742 has eight memory channels where a desktop has two. The eight-core rows move closer to a desktop's CPU but keep the server's memory. Strata's own table for an RTX 5070 with 64 GB of RAM lists 79 tokens per second for IQ2_XS (measured, its details page says, with a fine-tune of the same size that runs at the same speed). That was measured by its author on different hardware, and I did not reproduce it.
What to take from this
- If you already run this model in llama.cpp and the whole model fits in your GPU, no score difference showed up in my tests. Strata was still faster at reading and writing in my runs.
- If your card is around 12 GB, this is the case Strata was built for, and the speed gap was large. I found no quality cost on these tests.
- If you compare runs or cache results, set the reproducibility switches, or treat a changed answer at temperature 0 as possible noise rather than a regression.
- The model's license is the Qwen Community License 1.0. The quantized repository lists Apache 2.0 in its card, but that does not change the base model's license.
What this does not show
One file (IQ2_XS), thinking off, and short answers of up to 1,024 tokens. Long reasoning and agent work, where small numerical differences have more tokens to grow, were not tested. Three hundred GSM8K questions, 164 HumanEval problems and 100 images can only reveal large gaps. Condition B for the score tests was added after I saw condition A. The speeds come from an A100 and a server CPU, timed by the client, with Strata drafting and llama.cpp not. Strata's log does not record how many threads it started in the eight-core run, so I could not check that. I did not test an SSD, other quantizations, or other llama.cpp offload strategies. Which Qwen file fits which card in plain llama.cpp is in How to Run Qwen Locally.
If you want the next measurement in your inbox, the subscribe box is just below.
Setup: one A100 80GB PCIe (driver 580.178.04), AMD EPYC 7742 (64 cores), 251 GB RAM, model files on a rotational disk. Model: ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF, IQ2_XS/ both shards (SHA-256 92cee27a… and 316b46f3…, matching Hugging Face) and mmproj-Qwen3.8-Flash-Next-BF16.gguf. Strata 6f32ec0 (engine 0.1.39) in its Docker image built with CUDA_ARCHITECTURES=80, set up with setup.py --setup --yes --model IQ2_XS --gguf-dir … --context 65536 --vision yes. llama.cpp d89651a, llama-server -m … --mmproj … -c 65536 -fa on -lm mmap -ngl 99 (--lazy-mode on for condition A scores, off elsewhere; --n-cpu-moe 42 for condition B). Requests: OpenAI chat completions, temperature 0, chat_template_kwargs: {"enable_thinking": false}, max_tokens 1024 (64 for images), one at a time. GSM8K: test set shuffled with seed 0, first 300, last number compared with the answer. HumanEval: all 164, the returned function run against the tests with a 10-second limit. COCO: val2017, 100 image/category pairs with exactly one instance covering at least 1% of the image, seed 0, one pair per image; miss = distance from the box centre in original pixels. Score tests used the two cards at the same time, one engine each; speed tests ran one at a time on one card with the other idle and a one-minute load average below 8. Eight-core rows: taskset -c 0-7 … -t 8 for llama.cpp, --cpuset-cpus 0-7 for Strata's container. VRAM: nvidia-smi every 0.5 s, peak. The plan committed before the first run, scripts (scripts/strata-repro/), raw responses and logs are in drafts/strata-repro/. Measured 2026-10-05.
New posts on this topic go into the weekly newsletter. One email a week at most.
Courses that go with this post
SOTAAZ courseCourses are sold one at a time. If you want all 13, there is a $199 lifetime bundle.
Subscribe to Newsletter
Related Posts

llama.cpp KV Cache Quantization: Which -ctk and -ctv to Use, Measured on the Current Build
Keep f16 if the KV cache fits. If not, set -ctk q8_0 -ctv q8_0: it saved 4.2 GiB at 64K but decoded at 56% of f16's speed there.

How to Run Qwen Locally: Which File Fits Your GPU (8, 12, 16 or 24 GB), Measured
Which Qwen GGUF fits an 8, 12, 16 or 24 GB GPU, with measured memory: 9B Q4_K_M 5.8 GiB, 27B Q3_K_XL 13.1 GiB, 35B-A3B 7.3 GiB with 30 layers' experts on the CPU.

DiffusionGemma on vLLM's Example Server: One Sentence of Context Took It From 54.5% to 31.8%
I ran the example server from the merged vLLM pull request that turns DiffusionGemma into a decision model on 154 BANKING77 messages. It caps a question at 26 options, so I split the 77 intents into nine groups and read twice: 54.5% at 119 ms with one noise draw, 51.3% with the default draws. My transformers version of the same two-stage read scored 68.2%, and one added sentence describing the input dropped the server from 54.5% to 31.8%.

