Where Decision Models Go Wrong: Four Situations Behind the Sentences Most of Them Miss
We re-read the per-sentence results of up to 14 Jev-style decision models on six datasets and kept the 433 sentences at least half of them got wrong. Grouping those sentences and checking a few simple features showed four recurring situations: the question's shape and the dataset's rule disagree, categories overlap, the sentence contains another option's name, and no option fits. Examples, a full table of groups, and what to check in your own options.

Where Decision Models Go Wrong: Four Situations Behind the Sentences Most of Them Miss
Over the past few weeks we measured a dozen Jev-style decision models (what Jev is) on six public datasets. Those posts asked which model scores higher. This one asks a different question. When most of the models get a sentence wrong, what kind of sentence is it?
That question is more useful when you design your own options. A sentence that trips up ten different models is unlikely to be one model's quirk, so it tells you something about the options you wrote or the inputs you send.
Nothing was rerun for this post. We re-read the per-sentence results already saved for the earlier posts and committed the grouping rules before counting anything. There are no model scores or rankings here.
What we used
- Datasets: TREC (500 questions, 6 answer types), AG News (200 articles, 4 topics), financial news tweets (200, 20 topics), MASSIVE (175 requests, 60 intents), CLINC150 (400 requests, 150 intents) and a 154-message BANKING77 sample (77 intents).
- Models: Jev 1.13 through its API; Kev 0.8B, 4B and 9B; Jeff v1.1 0.8B and 2B; laya; MoJev; NanoJev; openjev; the DiffusionGemma example server; and Nimble 9B, Tev1 4B and Tev1 0.8B in Ollama. Each dataset uses the models that were run on it: 14 on TREC and AG News, 11 on the tweets and BANKING77, 8 on MASSIVE and CLINC150.
- Condition: option names only, with no written definitions, as run in Kev vs MoJev, Kev on unseen tasks, Kev vs Jev, Jeff vs Jev, the Ollama post and the open alternatives post.
- Hard sentence: one that at least half the models on that dataset got wrong. There are 433 across the six datasets.
- Group: hard sentences with the same right answer and the same most common wrong answer. Every group with three or more sentences is in the table at the end. The sections below discuss some of them. The rule decides which sentences belong to a group. The headings are our reading of the sentences.
- Features: we also flagged sentences by simple features, decided in advance, and compared how many models got flagged and unflagged sentences wrong. Section 3 comes from that comparison, not from groups.
| Dataset | Options | Models | Average share of models wrong | Hard sentences | Wrong for every model |
|---|---|---|---|---|---|
| TREC | 6 | 14 | 28% | 87 of 500 | 5 |
| AG News | 4 | 14 | 16% | 21 of 200 | 4 |
| Financial tweets | 20 | 11 | 45% | 80 of 200 | 16 |
| MASSIVE | 60 | 8 | 33% | 45 of 175 | 11 |
| CLINC150 | 150 | 8 | 44% | 154 of 400 | 105 |
| BANKING77 | 77 | 11 | 38% | 46 of 154 | 13 |
CLINC150's 105 includes 100 out-of-scope requests that no option could answer. Section 4 covers them.
1. The question's shape points one way, the dataset's rule another
TREC labels a question by the kind of answer it expects. Many models read the question's surface instead. Three groups hold 73 of TREC's 87 hard sentences:
- Asked about a place, read as a thing (33 sentences). "What is the brightest star?" is
locationin TREC. 12 of 14 models got it wrong, 11 of them withentity. "What is the longest suspension bridge in the U.S.?" was wrong for 11, 9 of thementity. - Asked about a thing, read as a definition (22). "What is foot and mouth disease?" is
entitybecause the answer is a disease. All 14 models got it wrong, and 13 answereddescription, which is what "What is X?" usually asks for. - Asked about an organization, read as a thing (18). "George Bush purchased a small interest in which baseball team?" is
human, because TREC counts organizations as people. All 14 got it wrong, 13 withentity.
We had runs with TREC's one-line option definitions for all 14 models, so we compared afterwards, outside the pre-registered rules. Only one of these rules is written in the definitions: "human: a person, a group of people or an organization." The location definition lists cities, countries, mountains and states, not stars, and the entity definition does not mention diseases. All three groups still improved. The share of models wrong fell from 0.63 to 0.38 for the place group, 0.61 to 0.30 for the definition group and 0.69 to 0.41 for the organization group. The fourth TREC group went the other way. Its three sentences, such as "What is the sales tax in Minnesota?" (labelled entity; 13 of 14 models wrong without definitions, 12 of them with numeric), went from 0.86 to 1.00, so every model got them wrong with definitions. Across all 500 questions the share fell from 0.28 to 0.22. The groups were picked because they were hard without definitions, so part of any drop would happen on a rerun anyway.
2. Categories that overlap
Here the models' answers are often defensible, and the trouble is in where the dataset draws its lines.
- AG News, technology vs business (two groups, 7 and 5 sentences). "I.B.M. Agrees to Settle Part of Giant Pension Case" is labelled Sci/Tech, and 12 of 14 models said Business. "Halo 2 scores record sales of $125 million in first 24 hours" is labelled Business. 10 of 14 got it wrong, 9 of them with Sci/Tech. A tech company's money story fits both.
- Financial tweets, sector vs report type (8). "$MLI - Mueller GAAP EPS of $3.65, revenue of $1.15B" is labelled Financials, and 9 of 11 models said Earnings.
- MASSIVE, date vs calendar (3). "what day of the week does christmas fall on this year" is labelled as a date question. 5 of 8 models got it wrong, 4 of them filing it under calendar. This group is small and only about half the models disagreed.
The tweets also had sentences with almost nothing in them. "$AMTB $ALSN $VZ $BAESY $VLTA" is five stock tickers, labelled Stock Commentary. 10 of 11 models got it wrong, and their answers scattered: 5 Stock Movement, 4 Financials, 1 Markets. With no content to read, nothing holds the answer in place.
3. The sentence contains another option's name
We flagged a sentence when it contained a word from some other option's name (four letters or longer) and no word from the right option's name, then compared flagged and unflagged sentences.
- BANKING77: flagged sentences had 25.7 points more of the models wrong (95% interval 13.5 to 37.9, 32 sentences). "How can I use American Express to add money to my account?" is labelled
supported_cards_and_currencies. All 11 models got it wrong, and 8 of them answeredtopping_up_by_card. - CLINC150: the pre-registered comparison showed +24.5 points. That figure includes 59 of the 100 unanswerable requests. Their label
ooshas no qualifying word, so any of them with another option's word got flagged. With all 100 removed, which we decided after seeing the result, the gap is 6.7 points (1.2 to 12.2, 80 sentences). - AG News and the tweets: we can't tell (intervals cross zero). TREC and MASSIVE had too few flagged sentences to judge.
This is a correlation. It tells you where to look, not that renaming an option fixes it. How much option names outweigh definitions is measured separately for laya in our label-bias rerun.
Some of these "errors" look as reasonable as the labels. "Why did the app refuse to make an approved payment" is labelled reverted_card_payment?, where the question mark is part of BANKING77's label name. All 11 models got it wrong, and 8 of them said declined_card_payment. When most models pile onto the same other answer, it is worth checking the label before blaming the models.
Two other features we registered gave narrower results. Negation (not, no, n't) went with more errors on BANKING77 only (+11.4 points, interval 0.1 to 23.0, 35 sentences), and we couldn't tell on the three other datasets with enough such sentences. Questions of five words or fewer had fewer errors on TREC (−15.3 points, −18.3 to −12.3), and sentences with digits had fewer on CLINC150 (−13.6 points). None of these holds across datasets, so we don't generalize them.

4. None of the options fits
CLINC150 contains 100 requests outside all 150 intents, such as "is the earth flat" or "watch the fbi." In the earlier Kev vs Jev runs we deliberately gave no "none of these" option, to see whether the probabilities would flag them. Every model therefore had to pick some intent, and the answer was always wrong.
Where they landed is the useful part. Of 800 answers (100 requests × 8 models), 128 went to no and 49 to yes. Some of these requests are yes/no questions (about a fifth start with is, can, do and the like), and the models answered them instead of classifying them: "is the earth flat" got no from 4 of 8 models. The rest spread across specific intents: 74 to reminder_update, 67 to what_can_i_ask_you, 55 to accept_reservations. The fix is either an explicit "none of these" option or a probability threshold that hands low-confidence answers to a person. The Kev vs Jev post found the threshold workable on this set: Jev's top probability separated these requests from in-scope ones at AUROC 0.94.
What to check
- Is the rule visible from the names? If organizations count as
human, a model reading only names won't know. On TREC, definitions helped three of the four groups and made one worse, so measure before and after. - Which pairs do many models confuse? Run a few models on your evaluation set and list the right-answer and wrong-answer pairs they share. Merge those categories, or define the boundary in one sentence each.
- Do your option names share words with typical messages? On BANKING77 that went with many more errors. Check for such pairs first; we did not test whether renaming helps.
- Is there a way out? Without "none of these" or a probability threshold, yes/no questions get answered and other off-topic requests land on specific intents.
- Do most models agree on a different answer? Check the label. Some are mislabelled or ambiguous.
- Does your evaluation set include these? Short "What is X" questions, tech-company money news, ticker-only posts, and requests that fit nothing.
Limits
- Option names only. With definitions, some groups shrink and one grew, as the TREC check shows. We did not repeat the whole analysis with definitions.
- English only, and samples of 154 to 500 sentences per dataset. The model set differs by dataset.
- The models are not independent. Several share bases or training data, so "10 of 11 models" is weaker evidence than ten unrelated judges would be.
- BANKING77's errors scatter across 77 intents, so no group reached three sentences there.
- The other-option-word flag is crude word matching. It misses synonyms and counts coincidences.
- Group headings are our reading of the sentences. The groups themselves come from the rule.
All groups with three or more hard sentences
| Dataset | Right answer → most common wrong answer | Sentences | Discussed |
|---|---|---|---|
| TREC | location → entity | 33 | §1 |
| TREC | entity → description | 22 | §1 |
| TREC | human → entity | 18 | §1 |
| TREC | entity → numeric | 3 | §1 |
| AG News | Sci/Tech → Business | 7 | §2 |
| AG News | Business → Sci/Tech | 5 | §2 |
| AG News | World → Business | 3 | |
| Tweets | Financials → Earnings | 8 | §2 |
| Tweets | Fed / Central Banks → Macro | 3 | |
| Tweets | Markets → Stock Movement | 3 | |
| Tweets | Stock Commentary → Stock Movement | 3 | §2 |
| Tweets | Stock Commentary → Company / Product News | 3 | |
| Tweets | Treasuries / Corporate Debt → Financials | 3 | |
| MASSIVE | datetime_query → calendar_query | 3 | §2 |
| CLINC150 | oos → no / yes / fun_fact / accept_reservations / time / smart_home / directions / what_can_i_ask_you | 26 / 9 / 8 / 5 / 4 / 3 / 3 / 3 | §4 |
The plan (committed before any counting), the analysis script, the selected result files and the after-the-fact checks are in drafts/failure-clusters/. The newsletter covers the next post when it is published.
