Models & Algorithms•SOTAAZ Lab••KR

Does laya Read Your Definitions or Your Labels? A Pre-Registered Rerun of the Label-Bias Test

We reran the 'labels override definitions' tests from arXiv 2610.02586 on two laya 0.3.7 checkpoints with our own TREC and AG News sets. Changing one line of prompt rendering made predictions identical whatever the labels were called. When labels contradicted the definitions, laya's TREC accuracy fell from 88.6% to 22.2%, and its top probability ranked right and wrong answers backwards (AUROC 0.36), the same direction the paper reports.

Does laya Read Your Definitions or Your Labels? A Pre-Registered Rerun of the Label-Bias Test

Does laya Read Your Definitions or Your Labels? A Pre-Registered Rerun of the Label-Bias Test

A typed decision model takes an input and a fixed question, and returns a probability for each option you defined. Each option has a short label and a written definition, and the definition is where you state the rule you want applied. A paper posted on October 1 asks whether the open implementations actually follow that rule.

Azizi, Baghaei Potraghloo and Pedram (arXiv 2610.02586) say they mostly don't. From the abstract: "Deleting every definition leaves accuracy unchanged (laya-td: 0.8559 against 0.8487)," and "laya writes each option as '{label}: {definition}', while von writes only the definition. Changing that expression in both directions, with no weight changed, makes all three laya checkpoints exactly invariant." The paper says its code, benchmark generator and per-item outputs "are released," but we found no link to them on the abstract page or anywhere in the PDF.

We had a reason to check. In our comparison of open Jev alternatives on September 23, adding one-line definitions to TREC moved laya by only 2 points (433 to 443 of 500, p = 0.25), while the same definitions moved a DiffusionGemma server by 20 points. That looked like the same effect. So we reran the paper's tests on our own data, with the decision rules and the sentences we would write for each outcome committed before any run.

What we ran

  • Models: two of the paper's three laya checkpoints, from laya 0.3.7 at their shipped settings. One is the default checkpoint convaiinnovations/laya, which the paper calls LAYA-EN and we call laya below. The other is the typed-decisions checkpoint, the paper's LAYA-TD and our laya-td. We did not test the multilingual checkpoint, von or the Qwen2.5 readouts.
  • Data: the same files as the September post. TREC, 500 questions and 6 answer types, is the main set. AG News, 200 articles and 4 topics, is secondary, because laya's documentation lists it as training data.
  • The line in question: in laya 0.3.7, render_options in laya/common.py builds each choice option as "%s: %s" % (label, definition). Our patch swaps that function, inside the runner only, for one that returns the definition alone. The library files are untouched.
  • Scoring: always against the category the definition describes, even when we renamed or scrambled the labels.

Seven conditions, each over the same items:

ConditionOptions passedRendering
C0real label + definitionas shipped (label: definition)
C1real label onlyas shipped
C2real label + definitionpatched (definition only)
C3A, B, C… + definitionas shipped
C4A, B, C… + definitionpatched
C5shifted labels + definitionas shipped
C6shifted labels + definitionpatched

"Shifted" means each definition gets the next category's label: TREC's "a person, a group of people or an organization" is shown as location, "a definition, explanation, reason or manner" as human, and so on around the list. A model that follows labels lands exactly one category off.

Comparisons are paired on the same items: McNemar's exact test at p < 0.05, with a 10,000-sample bootstrap interval for the accuracy difference. We ran C0 twice to check determinism, and both runs matched to the last probability. The C0 and C1 counts, 443 and 433, also match the September post exactly.

Horizontal bar chart of TREC accuracy for laya (blue) and laya typed-decisions (orange) across five rows: label plus definition 88.6% and 87.8%, label only 86.6% and 81.8%, A to F labels plus definition 88.2% and 88.2%, labels shifted to contradict the definitions 22.2% and 34.2%, definition-only rendering 91.2% and 89.0%

Deleting the definitions: one checkpoint shrugs, the other doesn't

laya went from 443 to 433 correct on TREC without definitions (−2.0 points, interval −5.0 to +1.0, p = 0.25). At 500 questions we can't call that a difference, which matches the paper's "unchanged."

laya-td is a different story. It dropped from 439 to 409 (−6.0 points, interval −9.6 to −2.4, p = 0.002), so on TREC the definitions did help this checkpoint. The paper's laya-td number comes from its own task suite, and we are not saying it is wrong. On our task, "unchanged" held for one checkpoint and not the other.

On AG News neither checkpoint moved by more than 2 of 200, which we can't distinguish from zero.

Renaming options to A, B, C…: no detectable change on TREC

The paper reports +0.1511 accuracy from renaming options to A and B. That number comes from PolicyBench, a synthetic routing suite the authors built so the rule appears only in the definitions. On classification tasks, the paper's mitigation section says renaming "costs accuracy instead, because there the class name is itself a good predictor of the class." TREC is that kind of task: human and location are good clues on their own.

So we expected renaming to hurt here, and we couldn't detect that it did. laya scored 441 against 443 and laya-td 441 against 439, and neither difference passes the test (p = 0.83 and 0.84). AG News leaned the other way. Every item that changed was a loss: laya lost 4 and laya-td lost 3 of 200, with none gained. McNemar can't call that at this size (p = 0.125 and 0.25). The bootstrap interval excludes zero for laya (−4.0 to −0.5 points) and touches it for laya-td (−3.5 to 0.0).

The one-line patch: identical predictions, whatever the labels say

With definition-only rendering, C2 (real labels), C4 (A to F) and C6 (shifted labels) produced the same prediction on every item, in all four model and dataset combinations. The largest difference in any option's probability was 0.0. Once the label is no longer written into the input, its name has no path to the model. That reproduces the paper's +0.0000.

When labels contradict definitions, laya follows the label

This is the result that matters if you ever rename an enum and forget a definition, or inherit labels that drifted from what they were meant to cover. With shifted labels, laya's TREC accuracy fell from 88.6% to 22.2% (111 of 500). In 333 of the 500 questions it picked the option carrying the right category's name, even though that option's definition described something else.

"Who was Galileo?" is a human question. Under shifted labels, laya chose the option labelled human whose definition reads "a definition, explanation, reason or manner," with probability 0.986.

laya-td held up somewhat better, at 34.2% with 260 label-following answers, but it is still far below its 87.8%. The patched rendering brings both back (91.2% and 89.0%). AG News shows the same pattern: 37 and 57 of 200 correct under shifted labels, with 156 and 139 answers following the label.

The paper's "misleading labels" condition, on its fixed-label classification tasks (SST-2, RTE and CB), drops LAYA-EN to 0.124 and LAYA-TD to 0.136 accuracy (Table 6). Our shifted labels are a different construction on different data, so the levels differ (22.2% and 34.2%), but the direction is the same.

The top probability ranked right and wrong answers backwards

A common safeguard is to route anything below some confidence to a person, which we looked at in Which Mistakes Cost You. That safeguard only works if high confidence goes with being right. We measured this as AUROC: how well the top probability separates correct answers from wrong ones, where 0.5 is a coin flip.

With labels as shipped, laya's TREC AUROC was 0.89 (interval 0.84 to 0.93). At a 0.9 threshold, laya accepted 314 answers on its own and 98.1% of them were right, while the 186 it handed to a person were 72.6% right. That is the safeguard working.

With shifted labels the AUROC was 0.36 (0.31 to 0.42). The whole interval sits below 0.5, so the top probability now ranks errors above correct answers. At the same 0.9 threshold, laya accepted 115 answers and only 4.3% of them were right. It handed 385 to a person, and 27.5% of those were right. The threshold kept the mistakes and passed along the better answers.

For laya-td on TREC the shifted-label AUROC was 0.52 (0.47 to 0.57), so we can't say which direction it ranks. On AG News both checkpoints ranked backwards: 0.22 and 0.26, with both intervals below 0.5.

This reproduces the direction of the paper's §5.8, which reports the same reversal: 0.194 for LAYA-EN and 0.243 for LAYA-TD under misleading labels (Table 6). Our one disagreement is laya-td on TREC, where we can't call a direction. The paper used laya's separate confidence field, which rarely equals the top probability. Recomputing ours with that field, also after the fact, gives the same conclusions: 0.35 for laya on TREC, 0.52 for laya-td on TREC, and 0.29 and 0.30 on AG News.

laya prints a warning when it loads either checkpoint: one shipped temperature is outside its accepted range and gets clamped to 0.5. That temperature belongs to questions with 11 or more options. Our tasks have 6 and 4 options, so it was not used in any of these runs.

One result we did not pre-register

Comparing C2 with C0 (patched against shipped rendering, real labels) was not one of our registered questions, so read this as a post hoc observation. On TREC, laya scored 456 against 443 (+2.6 points, p = 0.019). laya-td gained 6 (p = 0.42), and on AG News both moved by 1 to 3 items in either direction. That is one checkpoint on one dataset, so it is not a recommendation to patch for accuracy.

What to take from this

  • On real tasks with sensible labels, the label dependence cost little on TREC. We couldn't detect a change from renaming options, and deleting definitions cost laya-td but not laya.
  • The risk is mismatch. If a label and its definition disagree, laya follows the label, and its top probability can point the wrong way. That combination is hard to spot, because nothing errors and the probabilities look plausible.
  • Check the probabilities, not the predictions. The paper's two-call test (§5.6) sends the same question twice, changing only the labels, and compares the returned distributions. If they differ at all, the model is reading the labels. No gold answers are needed. Counting changed predictions is not enough. In a comparison we added after seeing the results, not one we pre-registered: on TREC, swapping in A to F changed laya's prediction on only 27 of 500 questions, which looks harmless. Its distribution moved by more than 0.01 (total variation) on 359 of them. That is the same model that drops to 22.2% when the labels contradict the definitions.
  • If the test says labels matter, there are two fixes. If you control the rendering, writing only the definition removes the dependence entirely. Without touching the library, you can name the options A, B, C… and put the whole rule in the definitions. On TREC that cost nothing we could detect (441 vs 443). The paper lists it as its first mitigation but reports a cost on classification tasks, and our AG News run lost a few items, so measure it on your own task.

Limits

  • Two laya checkpoints, the paper's LAYA-EN and LAYA-TD. We did not test LAYA-ML, von or the Qwen2.5 readouts, and we did not test hosted Jev. The paper doesn't measure hosted Jev either.
  • Two public classification sets. AG News is in laya's training data, which is why TREC carries the conclusions.
  • Shifted labels are a deliberately harsh contradiction. Real drift is usually partial, such as one renamed option or a definition that grew while its label stayed the same. We didn't measure partial drift.
  • Each condition ran once. Inference is deterministic, and our repeated C0 run was identical, so seeds don't apply here.

The plan, runner, raw per-item results and analysis are in drafts/label-bias/ (plan committed in 925bfb4, before the first run). The newsletter covers the next rerun when it is published.