Models & Algorithms•SOTAAZ Lab••KR

How Far Can You Shrink a Decision Model? decider-4b from BF16 to Q2_K on llama.cpp

We ran the open decision model decider-4b at five GGUF sizes, from BF16 (10.4 GiB of GPU memory) to Q2_K (4.4 GiB), on 500 TREC questions. Accuracy differences were too small to detect at every size. The probabilities moved: log-loss got slightly better at Q4_K_M and worse at Q8_0 (barely), Q3_K_M and Q2_K. Q2_K changed 13% of the answers and Q3_K_M cut the share of answers above 0.9 from 40% to 23%. Q4_K_M halved the memory; no accuracy difference was detected and its log-loss improved.

How Far Can You Shrink a Decision Model? decider-4b from BF16 to Q2_K on llama.cpp

How Far Can You Shrink a Decision Model? decider-4b from BF16 to Q2_K on llama.cpp

A decision model reads some content and a question with fixed answers, and returns a probability for each answer. What you are buying is the probability: you let the application act when it is high and send the case to a person when it is not. So when you quantize one to fit a smaller GPU, the question is not only whether it still picks the right answer, but whether its probabilities still mean the same thing.

decider-4b is an open 4B decision model with official GGUF files for llama.cpp. Its author already published accuracy, log-loss and calibration for BF16, Q8_0 and Q4_K_M on a large regression set: Q4_K_M agreed with the bf16 weights on 97.06% of 144,226 rows (GGUF card). Two things were missing for someone choosing a file: the sizes below Q4, and how much memory each size actually saves. We measured both, on a task outside the author's published training list, and with the same "labels that contradict their definitions" test we used in our post on Decision-1.

The finding we would lead with is not the file size. No accuracy difference was detected at any size, but under the same 0.9 rule, Q3_K_M accepted 23.2% of the answers against 40.2% for BF16, so a person would have to handle 77% of the questions instead of 60%. Quantization changed the workload, not just the memory.

How we measured

  • Files: the official BF16, Q8_0 and Q4_K_M GGUFs of decider-4b v2.1 (checksums match the repository), plus Q3_K_M and Q2_K that we made from the official BF16 file with llama.cpp 4da6337's llama-quantize at default settings, without an importance matrix. Those two are ours, not the author's.
  • Runtime: decider-ai 1.6.0 on llama-cpp-python 0.3.36 (CUDA), all layers on one A100, context 4,096 for every size. GPU memory is what the process used after loading the model and answering the first question.
  • Data: 500 TREC questions (six answer types), each option written as its name and definition. In the decider-ai 1.6.0 source (decider/data/core.py), decider's public task configuration marks trec as held-out and ag_news, which we use once below, as a training task. That says what decider's own fine-tuning used; it doesn't mean the base model it was built on never saw TREC.
  • Main measures, fixed before scoring: accuracy and log-loss (NLL) against BF16, for each of the four smaller files. That is eight comparisons, and their p-values are corrected together with Holm's method. A difference counts only if it survives the correction, and we report its direction either way. The 95% intervals we show are ordinary bootstrap intervals, not adjusted for multiple comparisons.
  • Also reported, without a verdict: calibration error (ECE), how well the top probability separates right from wrong (AUROC), how many answers clear 0.9, and how often each file's answer differs from BF16's.

Speed is not measured: the GPU was shared with other work.

Results

Three panels across five files: BF16, Q8_0, Q4_K_M, Q3_K_M (ours), Q2_K (ours). GPU memory: 10.4, 6.8, 5.1, 4.7, 4.4 GiB. Change in log-loss against BF16, with 95% intervals: Q8_0 +0.002, Q4_K_M −0.022, Q3_K_M +0.037, Q2_K +0.060. Answers different from BF16: 0.2%, 1.2%, 3.2%, 13.2%.
FileSize on diskGPU memoryMemory saved vs BF16Accuracy (change, 95% interval)Log-loss (change, 95% interval)Answers different from BF16
BF168.42 GB10.4 GiB—86.4%0.387—
Q8_04.48 GB6.8 GiB3.7 GiB (35%)86.2% (−0.2, −0.6 to 0.0)0.389 (+0.002, +0.001 to +0.003)1 (0.2%)
Q4_K_M2.71 GB5.1 GiB5.3 GiB (51%)86.4% (0.0, −1.0 to +1.0)0.365 (−0.022, −0.031 to −0.015)6 (1.2%)
Q3_K_M (ours)2.26 GB4.7 GiB5.8 GiB (55%)86.0% (−0.4, −2.0 to +1.2)0.425 (+0.037, +0.022 to +0.051)16 (3.2%)
Q2_K (ours)1.92 GB4.4 GiB6.1 GiB (58%)85.6% (−0.8, −3.8 to +2.2)0.447 (+0.060, +0.021 to +0.099)66 (13.2%)

Disk sizes are in GB (10⁹ bytes), GPU memory in GiB (2³⁰ bytes). Accuracy changes are in percentage points.

Accuracy: no detectable difference at any size. The changes against BF16 were −0.2, 0.0, −0.4 and −0.8 points. Even for Q2_K the 95% interval runs from −3.8 to +2.2 points, so 500 questions can't rule out a loss of a few points.

Log-loss moved, in both directions. Verdicts after the Holm correction of the p-values (intervals unadjusted):

  • Q8_0 was worse by 0.0017 (interval 0.0006 to 0.0028). Real, but tiny.
  • Q4_K_M was better by 0.022 (interval −0.031 to −0.015).
  • Q3_K_M was worse by 0.037 (0.022 to 0.051).
  • Q2_K was worse by 0.060 (0.021 to 0.099).

Q4_K_M coming out better than BF16 is a measurement on 500 questions, not a reason to expect quantization to help. The author's own table also shows Q4_K_M log-loss slightly worse on in-task rows and slightly better on held-out rows.

Different answers are not the same as wrong answers. Q2_K changed 66 answers. In 28 of them Q2_K was right and BF16 wrong, in 32 the reverse, and in 6 both were wrong. That's why accuracy barely moved while the answers changed a lot. For a pipeline that has been tuned on BF16's behavior, 13% of decisions changing is a change, whatever it does to accuracy. At Q4_K_M, 6 answers changed: 3 each way. Our agreement figures (99.8% for Q8_0, 98.8% for Q4_K_M) are on a different and much smaller set than the author's 99.31% and 97.06%, so we don't compare them.

What happens to "act when above 0.9"

FileAnswers at or above 0.9Wrong among them (count, 95% interval)ECE
BF1640.2%3 of 201 (0.5% to 4.3%)0.054
Q8_039.4%3 of 197 (0.5% to 4.4%)0.049
Q4_K_M39.8%3 of 199 (0.5% to 4.3%)0.070
Q3_K_M23.2%0 of 116 (0.0% to 3.2%)0.086
Q2_K31.0%2 of 155 (0.4% to 4.6%)0.098

Intervals are Wilson 95% intervals.

At Q3_K_M, 23.2% of answers reached 0.9, against 40.2% for BF16 (difference −17.0 points, interval −20.4 to −13.6). Among the 116 answers it accepted we observed no errors, which is not the same as showing it is safe: with 116 answers the interval still reaches 3.2%. The clear cost is that a 0.9 rule would send 77% of the questions to a person instead of 60%. If you swap files under a fixed threshold, the workload changes even when accuracy doesn't.

ECE against BF16: +0.032 at Q3_K_M (interval 0.006 to 0.058) and +0.044 at Q2_K (0.014 to 0.068); at Q8_0 (−0.006) and Q4_K_M (+0.016) the intervals include zero. These are reported numbers, not verdicts: the pre-registered verdicts cover accuracy and log-loss only.

With labels that contradict their definitions

We also gave every file the swapped options from the Decision-1 post: each definition labeled with another category's name, the correct answer always being the one the definition describes.

FileAccuracy, swappedFollowed the nameAUROC, swappedAnswers changed by reversing option order
BF1670.6%23.6%0.69911
Q8_071.0%23.2%0.69712
Q4_K_M77.4%16.4%0.76110
Q3_K_M68.8%23.6%0.60813
Q2_K66.2%23.6%0.63341

Every file followed the definitions more often than the names, and every file lost accuracy compared with normal labels (−9.0 to −19.4 points). The quantized files did not follow names more often than BF16. Every file changed some answers when the options were reversed, and Q2_K changed the most: 41 of 500, against 10 to 13 for the other files (11 for BF16). On the 200 AG News headlines, Q2_K's accuracy on swapped labels was 80.0% against 90.0% for BF16.

Which file to use

This part is our judgment from the tables above, not a pre-registered verdict.

  • Q4_K_M halved the GPU memory (10.4 to 5.1 GiB). On this task no accuracy difference was detected and its log-loss improved. 39.8% of its answers reached 0.9, against 40.2% for BF16 (−0.4 points, interval −2.4 to +1.6). It is the file we would start from.
  • Q8_0 saves less memory and showed no advantage here over Q4_K_M.
  • Q3_K_M and Q2_K save another 0.4 to 0.7 GiB. In exchange, log-loss got worse, Q3_K_M put far fewer answers above 0.9, and Q2_K changed one answer in eight; reversing the option order changed 41 of its answers against 11 for BF16. If memory is that tight, re-check your threshold on your own labeled cases before switching.

Limits

  • One task, 500 questions, one run per file (repeated runs gave identical probabilities on every file).
  • Our Q3_K_M and Q2_K are default llama-quantize output without an importance matrix; better-made low-bit files may behave differently.
  • GPU memory includes the CUDA context and depends on the context size (4,096 here).
  • TREC is a short, six-way task. Longer inputs and many-option questions may react differently to quantization.

The plan, scripts, file checksums and every answer are in the reproduction package below.

Files for this post

Reproduction package

The harness, raw logs and result CSVs behind this post. No account needed; the README has a command that recomputes the tables without a GPU.

decision-quant-repro.zip · 2,683 KB

Download