Can You Act on Microsoft-Decision-1's Probabilities? We Swapped the Labels on 500 Questions
We sent 500 TREC questions to Microsoft-Decision-1 and Jev 1.13 through OpenRouter, and ran two local decision models, decider-4b and laya, with normal labels and with option names swapped against their definitions. Decision-1 followed the definitions (95.2% against 96.4%, not a difference by our rule) and at a 0.9 threshold handled 80% of the swapped cases on its own, 5 of those 400 wrong. laya followed the labels and was wrong on 110 of the 115 swapped cases it would have handled. The same request sent twice gave slightly different probabilities on Decision-1.

Can You Act on Microsoft-Decision-1's Probabilities? We Swapped the Labels on 500 Questions
Decision models don't write text. You give them some content and a question with a fixed set of answers, and they return a probability for each answer. Microsoft's new Microsoft-Decision-1 is one of them, and the OpenRouter page describes its output as "a calibrated probability for each fixed answer option" that "can determine when an application acts, defers, or asks for review."
That is the use we care about: let the application act on its own when the probability is high, and send the rest to a person. Microsoft's post says "a 90% prediction should be right about nine times out of 10 on representative cases" and charts a calibration score over 36 benchmarks, where 100 means confidence matches how often the model is right: Decision-1 92.2, third of seven models, with Jev 1.13.0 at 93.7. The post doesn't say how the score is computed. It also reports that the decision changed on 1.3% of requests on average when the same request was altered in eight ways, and on none when option descriptions were paraphrased or the options reversed or shuffled. Those alterations all keep the meaning.
Earlier this week we found that laya, an open decision model, follows the option name over its definition: when the names were swapped against the definitions, its accuracy on TREC fell from 88.6% to 22.2%, and its probabilities ranked its errors above its correct answers. So we asked the same of Decision-1, and of two other decision models, before letting any of them act alone:
- With normal labels, how many answers would a 0.9 threshold let through, and how many of those are wrong?
- When option names and definitions disagree, which one does each model follow, and what happens to those numbers?
If you are setting up Decision-1 for the first time, our field guide covers the request format and what we ran into.
How we measured
- Data: 500 questions from TREC (what kind of answer a question asks for: abbreviation, entity, description, human, location or numeric), the same file and definitions as in our earlier post. A second set of 200 AG News headlines is in the reproduction package; we only mention it where it disagrees. In the decider-ai 1.6.0 source (
decider/data/core.py), decider's public task configuration markstrecas held-out andag_newsas a training task. That describes decider's own fine-tuning, not whether the base model it was built on ever saw TREC. - Models: Microsoft-Decision-1 and Jev 1.13 through OpenRouter's System One endpoint (responses carried the versions
microsoft-decision-1-20261009andjev-1.13-20260917), and two local models: decider-4b v2.1 (official BF16 GGUF on llama.cpp) and laya 0.3.7. - Conditions, one question per request:
- Normal: each option is its real name plus its definition.
- Swapped: every definition gets the next category's name, so "a person" is labeled location. The correct answer is always the category the definition describes.
- Reversed: the normal options in reverse order.
- Repeated: the normal request sent a second time.
- What "act alone" means here: the application accepts the answer when the model's top probability is at least 0.9. We report three numbers: how many answers it accepts, how many of the accepted are wrong, and how many wrong answers get through out of all 500. 0.9 is a common diagnostic point, not a recommendation. Note that Decision-1 and Jev also return a separate
confidencefield that is not the top probability; we used the top probability throughout. - Rules: we fixed the conditions, metrics and the sentence to write for each outcome before scoring, and committed them first. The API requests for each question were sent in shuffled order within a few seconds of each other. All 8,400 API requests succeeded (36 were rate-limited and resent). Total API cost: $0.094.
With normal labels
| Model | Accuracy | Accepted at 0.9 | Wrong among accepted (count, 95% interval) | Wrong that got through (of 500) |
|---|---|---|---|---|
| Decision-1 | 96.4% | 93.2% | 6 of 466 (0.6% to 2.8%) | 6 |
| Jev 1.13 | 93.4% | 84.0% | 14 of 420 (2.0% to 5.5%) | 14 |
| decider-4b (BF16) | 86.4% | 40.2% | 3 of 201 (0.5% to 4.3%) | 3 |
| laya | 88.6% | 62.8% | 6 of 314 (0.9% to 4.1%) | 6 |
Intervals are Wilson 95% intervals.
On this task Decision-1 let the most answers through at 0.9 and was wrong on 6 of the 466 it accepted. decider-4b was more cautious: it reached 0.9 on only 40% of questions, so a 0.9 rule would send 60% of the work to a person. The same threshold means very different workloads on different models.
The left panel below shows the whole trade-off: accept answers from most to least confident, and track how many of the accepted are wrong.

When the names and definitions disagree
| Model | Accuracy, swapped | Followed the definition | Followed the name | Accepted at 0.9 | Wrong among accepted (count, 95% interval) |
|---|---|---|---|---|---|
| Decision-1 | 95.2% | 95.2% | 0.6% | 80.0% | 5 of 400 (0.5% to 2.9%) |
| Jev 1.13 | 85.6% | 85.6% | 9.6% | 42.6% | 5 of 213 (1.0% to 5.4%) |
| decider-4b (BF16) | 70.6% | 70.6% | 23.6% | 12.4% | 1 of 62 (0.3% to 8.6%) |
| laya | 22.2% | 22.2% | 66.6% | 23.0% | 110 of 115 (90.2% to 98.1%) |
Decision-1 followed the definitions. Its accuracy went from 96.4% to 95.2%: 7 questions it had right became wrong and 1 the other way (p = 0.070), so by our rule we can't call that a difference, although the bootstrap interval for the change (−2.4 to −0.2 points) does not include zero. So we don't conclude that its performance held. What did change is how often it was sure: it accepted 80.0% at 0.9 instead of 93.2%, With normal labels 6 of the 466 accepted answers were wrong; with swapped labels, 5 of 400. Both round to about 1.3%. The probabilities still separated right from wrong answers (AUROC 0.869, interval 0.779 to 0.946).
Jev and decider-4b also followed the definitions more often than the names, but lost accuracy: Jev 7.8 points (p < 0.001), decider-4b 15.8 points (p < 0.001). Both became much less sure, so a 0.9 rule would have sent most swapped cases to a person (57% for Jev, 88% for decider-4b), and what they did accept was mostly right.
laya followed the names, as in our earlier test. The problem for automation is in the last column: of the 115 swapped questions where laya was at least 90% sure, 110 were wrong. Its probabilities ranked errors above correct answers (AUROC 0.362, interval 0.311 to 0.415). Under this condition, a 0.9 threshold on the top probability did not filter out the confident errors.
Same meaning, different order or a second try
| Model | Answers changed by reversing the options | Answers changed by sending the same request again |
|---|---|---|
| Decision-1 | 0 of 500 (95% upper bound 0.76%) | 1 of 500 |
| Jev 1.13 | 8 | 4 |
| decider-4b (BF16) | 11 | 0 |
| laya | 40 | 0 |
On our 500 TREC questions, reversing the options did not change any Decision-1 answer. That is consistent with Microsoft's report of no changes under reordering, on different data.
Sending the exact same request twice is where the two API models differ from the local ones. Decision-1 changed 1 answer, and its probabilities differed on 439 of 500 questions. Across all 500 the median difference was 0.0004, but 41 differed by more than 0.01 and the largest by 0.060. Jev changed 4 answers. The local models returned identical probabilities both times. If you tune a threshold on a hosted model, a probability near the threshold can land on either side on the next call.
On the 200 AG News headlines, three results differed from TREC:
- Decision-1 and option order: reversing the options changed 3 answers, and sending the same request again changed 2. With both above zero, we don't say whether order mattered there.
- decider-4b on swapped labels: accuracy barely moved (90.5% to 90.0%, not a difference by our rule), against a 15.8-point drop on TREC. AG News is a training task for decider-4b (see below), TREC is not.
- Jev on swapped labels: the AUROC interval included 0.5 (0.62), so we can't say whether its probabilities still ranked errors correctly.
What this does and doesn't show
- It shows that on this task, with a 0.9 rule on the top probability, Decision-1 would have handled most questions alone with about 1% of the accepted ones wrong, including when the option names contradicted their definitions.
- It shows that "high probability" is not safe by itself: laya was over 90% sure on 115 swapped questions and wrong on 110 of them.
- It does not show that Decision-1 is calibrated in general. TREC is a short, clean six-way task. Your labels, your data and your threshold decide the numbers; test them on a few hundred of your own labeled cases before letting the model act alone.
- It does not compare speed. We sent requests one at a time with waits for rate limits.
- The hosted model changes. These numbers are for
microsoft-decision-1-20261009; OpenRouter says the weights are updated continually.
Limits
- One main task (TREC, 500 questions) and one secondary (AG News, 200). Decision-1's and Jev's training data are not public, and TREC is a well-known dataset.
- One run per condition, except the repeated request.
- The 0.9 threshold and the top probability as the score are choices; with
confidenceinstead, the counts would differ. - laya's swapped, label-only and A/B results are reused from our earlier post after checking that a fresh normal run matched it exactly.
The plan, the scripts and every request and response (without account identifiers) are in the reproduction package below. How far you can shrink decider-4b before its probabilities change is in a separate post.
Files for this post
Reproduction package
The harness, raw logs and result CSVs behind this post. No account needed; the README has a command that recomputes the tables without a GPU.
decision-quant-repro.zip · 2,683 KB
