Models & Algorithms•SOTAAZ Lab••KR

How Much of an Uncertainty Score Is Just Answer Length? U-Space on Qwen3.5-4B

A new paper finds that answer length alone flags a language model's wrong answers almost as well as many uncertainty scores. We ran the authors' code on Qwen3.5-4B with their items: length alone reached AUROC 65.3 on three benchmarks, from 75.2 on TriviaQA to 48.6 on SuperGPQA. After controlling for length, U-Lens beat MSP (+8.4) and length alone (+4.4 before control), but its edge over mean log-loss (+1.1) was too small to call. On the three benchmarks in the main comparisons up to a third of answers hit the 16,384-token limit and were excluded; on Omni-MATH, left out of the comparisons, 92.5% did.

How Much of an Uncertainty Score Is Just Answer Length? U-Space on Qwen3.5-4B

How Much of an Uncertainty Score Is Just Answer Length? U-Space on Qwen3.5-4B

When a language model answers a hard question, you would like a number that says how likely the answer is to be wrong, so you can send the doubtful ones to a person. Many methods produce such a number from the model's own token probabilities. A paper from 6 October, U-Space (Braun et al.), asks an awkward question about them: how much of that number is just the length of the answer?

Reasoning models write longer when they struggle, and long answers are more often wrong. So a score that rises with length will look like it detects errors even if it knows nothing else. The paper measures this directly. Averaged over three 27-31B models, four benchmarks and three seeds, length alone separated right from wrong answers with AUROC 68.6, about as well as most uncertainty methods. Within narrow length bands its AUROC dropped by 16.2 points, close to chance. The authors' own method, U-Lens, which reads uncertainty from the model's hidden state, scored 71.1 and lost only 1.8 points under the same length control.

The models in the paper are large. We wanted to know whether the picture holds on a model people run locally, so we ran the authors' code on Qwen3.5-4B.

How we measured

  • Code and items: the authors' repository (s2labres/U-Space, commit cce4d26) without changes. The same four benchmarks (MMLU-Pro, Omni-MATH, SuperGPQA, TriviaQA) and the same 1,000 items each; the code checks the item lists against the paper's hashes, and all four matched.
  • Model: Qwen3.5-4B with the settings from the repository's own small-model config: thinking on, temperature 1.0, a 16,384-token generation limit, and the published 4B lens. We used one seed (42); the paper used three. We chose 4B because no lens of the kind the code needs is published for the 9B instruct model.
  • Scores: every method in the repository: U-Lens, its cone component A_cone, generation length alone, MSP, max entropy, mean NLL (mean log-loss of the generated tokens), Self-Certainty, DeepConf and predictive entropy. AUROC here is how well a score ranks wrong answers above right ones: 50 is chance, 100 is perfect. The length-controlled version is the authors': AUROC within ten length bands, averaged by band size.
  • Rules, fixed before running: three main comparisons, averaged over the benchmarks, each with a paired bootstrap interval. With three comparisons we use 98.33% intervals. A difference counts only if its interval excludes zero.

One thing went differently from the plan, and it matters for everything below. Qwen3.5-4B thinks long, and many answers ran into the 16,384-token limit. The authors' code excludes those, as it should, because a cut-off answer has no final answer to grade.

BenchmarkGraded (of 1,000)Cut off at the limit</think> more than onceAccuracy on graded
TriviaQA819180168.6%
MMLU-Pro951311882.2%
SuperGPQA6643072957.2%
Omni-MATH75925092.0%

The authors' code also excludes answers where the end-of-thinking marker appears more than once; each row adds up to 1,000.

On Omni-MATH only 75 answers finished, and they are the ones the model found easy enough to finish. Seventy-five answers are also too few for the authors' length-controlled AUROC (each of the ten length bands needs 20 answers). We decided before computing any AUROC to leave Omni-MATH out of the main comparisons and report it separately. On the other three, the graded answers are also the ones that finished within the limit, so the results describe "answers the 4B model completed", not all questions. The paper gave its 27B Qwen a 65,536-token limit.

Length alone is a strong signal, on some benchmarks

Two panels. Left: AUROC for spotting wrong answers on Qwen3.5-4B, averaged over TriviaQA, MMLU-Pro and SuperGPQA, before and after length matching. U-Lens 69.7 to 66.8, predictive entropy 69.7 to 66.4, mean NLL 69.7 to 65.7, DeepConf 67.4 to 64.9, Self-Certainty 66.7 to 64.9, max entropy 69.1 to 64.9, A_cone 61.5 to 60.4, MSP 55.8 to 58.4, generation length only 65.3 to 56.6. Right: median generated tokens for right and wrong answers. TriviaQA 2,000 and 5,073; MMLU-Pro 4,054 and 7,277; SuperGPQA 10,335 and 10,148.

Averaged over the three benchmarks, length alone reached AUROC 65.3, against the paper's 68.6 on larger models (different models and seeds, so we don't compare them as equal or different). But the average hides two very different benchmarks:

  • TriviaQA: 75.2. The median wrong answer was 5,073 tokens long and the median right answer 2,000. Length alone beat every uncertainty score on this benchmark, including U-Lens (71.5).
  • MMLU-Pro: 72.0. Wrong answers were longer here too (7,277 against 4,054 tokens).
  • SuperGPQA: 48.6, chance level. Right and wrong answers were about the same length (10,335 and 10,148 tokens), so length had nothing to offer.

So "length predicts errors" was true where the model rambled on questions it didn't know, and false where it wrote long answers to everything.

The three main comparisons

Comparison (mean over TriviaQA, MMLU-Pro, SuperGPQA)Difference in AUROC98.33% intervalVerdict
U-Lens vs length alone, no length control+4.4+1.0 to +7.8U-Lens separated wrong answers better
U-Lens vs MSP, length-matched+8.4+3.8 to +12.4U-Lens separated wrong answers better
U-Lens vs mean NLL, length-matched+1.1−0.5 to +2.6Too close to call with these items

These verdicts are on the three-benchmark average. On TriviaQA alone, length (75.2) was above U-Lens (71.5), as shown above.

On the 4B model, U-Lens beat length alone and MSP, as in the paper. The third comparison is where the 4B result differs: after length control, U-Lens and mean NLL (the simplest score built from the model's own probabilities) were within about a point, and the interval includes zero. In the paper's length-controlled table, U-Lens led mean NLL on every model, by 2.6 points on the 27B Qwen. We could not confirm that difference on 4B. With one seed, three benchmarks and the truncation above, that is "not confirmed here", not "wrong".

Every score lost ground under length control except MSP, which was near chance to begin with. Here is the full table next to the paper's averages, as context rather than a comparison:

Method4B raw4B length-matchedChangePaper raw (27-31B)Paper change
U-Lens69.766.8−2.971.1−1.8
Predictive entropy69.766.4−3.367.3−1.5
Mean NLL69.765.7−4.066.9−1.6
Max entropy69.164.9−4.267.7−5.0
DeepConf67.464.9−2.467.8−1.8
Self-Certainty66.764.9−1.766.9−4.4
A_cone61.560.4−1.1not in Table 1
MSP55.858.4+2.653.3−0.8
Length alone65.356.6−8.768.6−16.2

(4B: one seed, three benchmarks, truncated answers excluded. Paper: three models, four benchmarks, three seeds. The paper reports A_cone only length-matched, per model, in its Table 2. The paper's TokUR, Semantic Entropy and Feature-Gaps are not in the released code, so we didn't run them.)

Length alone kept 56.6 under length control, not 50: even inside one of ten length bands, longer answers were still a little more often wrong.

What this means for decision models

Our recent posts measured the same kind of AUROC on decision models such as Microsoft-Decision-1, which answer with a single option and its probability. A one-token answer has no length to leak, so this particular confound does not apply to those numbers. It applies when the model writes, and especially when it thinks out loud before answering.

What this does and doesn't show

  • It shows that on Qwen3.5-4B, with the authors' code and items, answer length alone flagged wrong answers well on TriviaQA and MMLU-Pro and not at all on SuperGPQA, and that U-Lens beat length alone and MSP on the three benchmarks combined.
  • It does not show that U-Lens is better than mean NLL on a 4B model, or that it isn't. The difference was too small to call here.
  • It does not describe questions the model never finished. Up to 31% of answers on SuperGPQA and 92.5% on Omni-MATH hit the token limit and were excluded. A cut-off answer is itself a strong sign of trouble, and none of the scores here were asked to detect it.

Limits

  • One model (Qwen3.5-4B), one seed, three benchmarks in the main comparisons.
  • A 16,384-token limit from the repository's small-model config, against 32,768-65,536 in the paper. A longer limit would grade more of the hard questions and could change every number.
  • The paper's ± values are spreads across models; ours are single runs.
  • We changed no code. The repository's requirements file omits accelerate, which we installed.

The plan, change log, our config and comparison script, every score file and the logs are in the reproduction package below. The model's full reasoning traces (about 30 MB) are not included; the code regenerates them.

Files for this post

Reproduction package

The harness, raw logs and result CSVs behind this post. No account needed; the README has a command that recomputes the tables without a GPU.

uncertainty-length-4b-repro.zip · 655 KB

Download

Stay Updated

Follow us for the latest posts and tutorials

Subscribe to Newsletter