Which Messages Should Go to a Person? A Classifier, a Decision Model and a Hand-Off, Measured
On BANKING77 a decision model after a classifier saved no hand-offs. On CLINC150 unknown questions broke the thresholds; adding 250 to validation halved the leaks.

Which Messages Should Go to a Person? A Classifier, a Decision Model and a Hand-Off, Measured
The last step of the text-classification guide is acting on a probability: answer the confident cases automatically and send the rest to a person. This post builds that system and measures it end to end. It has three stages:
- A small classifier (MiniLM embeddings with logistic regression, on a CPU; the same baseline as before) answers when its probability is above a threshold.
- Otherwise the classifier's top five candidates go to a decision model, which answers when its own probability is above a second threshold. I tried Nimble 9B and Tev1 4B locally through Ollama 0.35 (measured on their own here), and Jev 1.13 through its API, also with the full list of options.
- Otherwise a person. A message handed to a person is never counted as answered correctly.
The question is what the second stage buys: how many messages it takes off the person's desk at the same accuracy.
How the thresholds were chosen
Every setting was chosen on a validation set and scored once on a test set the selection never saw. For BANKING77, 77 intents of messages to a bank, I held out 10% of the training set (1,003 messages, stratified by intent) for validation and scored on the full test set of 3,080. For CLINC150, 150 intents for a virtual assistant, I used the official splits: validation 3,100 messages, test 5,500, and both contain out-of-scope questions that fit no intent. The thresholds are the ones that answered the most validation messages while keeping automatic answers at 95% (or 98%) correct. The test result is then reported as it came out, including when it falls short of the target.
The short answer
- On BANKING77, no decision model reduced the person's share. The messages the classifier was unsure about were just as hard for Nimble, Tev1 and Jev. At the 95% target the classifier alone answered 93.2% of the test set automatically, 96.2% of those correctly, and every cascade ended up with the same numbers or no better.
- On CLINC150, Jev and Nimble did help with the uncertain messages: at the 98% target Jev got 78.8% of them right where the classifier's first choice got 61.6% (p < 0.001); Tev1 was not clearly better.
- Unknown questions broke every threshold. CLINC150's validation set is 3% out-of-scope; its test set is 18%. Thresholds chosen for 95% accuracy produced 84.8% on test, with 592 of 1,000 out-of-scope questions answered as if they were in scope. Adding a decision model let more through.
- Putting unknown questions into validation cut the leaks by more than half. With 250 more out-of-scope messages in validation, leaks fell from 592 to 271, and Jev as the second stage then took 1.9 points of the traffic off the person without adding errors. The test still fell short of the 95% target.

BANKING77: the second stage had nothing to add
The classifier on its own was already strong: 92.9% of test messages right, and the correct intent was among its top five candidates for 99.3% of them. At the 95% target it answered 2,871 of 3,080 test messages automatically, 96.2% correctly (95% confidence interval 95.4–96.8%), and handed 209 to a person. At the 98% target it answered 82.9% at 98.0% (97.4–98.5%) and handed off 17.1%.
The second stage only sees the messages the classifier hands on, so the fair comparison is on those same messages: the classifier's first choice against each model's pick.
| BANKING77 test, messages the classifier handed on | 95% target (209) | 98% target (526) |
|---|---|---|
| Classifier's first choice | 47.4% | 67.9% |
| Nimble 9B, top 5 | 52.2% (p = 0.31) | 60.3% (p = 0.003) |
| Tev1 4B, top 5 | 51.2% (p = 0.46) | 54.9% (p < 0.001) |
| Jev 1.13, top 5 | 51.2% (p = 0.43) | 57.0% (p < 0.001) |
| Jev 1.13, all 77 options | 47.8% (p = 1.0) | 55.5% (p < 0.001) |
At the 95% target the models were no better than the classifier's own guess, and at the 98% target they were worse. Their probabilities did not separate right from wrong well enough either: on the validation messages the classifier handed on, Nimble's answers with a probability of 0.99 or more were right 58–82% of the time depending on the setting, out of only 10 to 30 such answers. So the threshold search on validation chose to accept no second-stage answers at all, for every model except Jev with all 77 options, which accepted answers at a probability of exactly 1.00. On test that added 35 automatic answers, a difference in automatic share of +0.3 points (−0.2 to +0.7) that is indistinguishable from none.
Two caveats sit with this table. Tev1 lists 3,000 BANKING77 examples in its training data, and even so it did not beat the classifier here. And the "accept nothing" choice was made on 83 validation messages at the 95% target and 195 at 98%, so it is a decision on a small sample.
CLINC150: the second stage helped on the uncertain messages
The same comparison on CLINC150 came out the other way.
| CLINC150 test, in-scope messages the classifier handed on | 95% target (48) | 98% target (349) |
|---|---|---|
| Classifier's first choice | 39.6% | 61.6% |
| Nimble 9B, top 5 | 66.7% (p = 0.011) | 69.6% (p = 0.018) |
| Tev1 4B, top 5 | 58.3% (p = 0.049) | 64.5% (p = 0.45) |
| Jev 1.13, top 5 | 68.8% (p = 0.007) | 78.8% (p < 0.001) |
So "the messages a classifier is unsure of are hard for everyone" held on BANKING77 and not on CLINC150. These runs do not say why. The obvious guess, that BANKING77's uncertain cases are near-duplicate intents, does not separate the two: in both datasets, when the classifier was wrong on a handed-on message, the right intent was usually its second choice (57–60% of the time on BANKING77, 63% on CLINC150 at the 98% target), and the most common mix-ups look alike, such as verify_my_identity against why_verify_identity in one and user_name against what_is_your_name in the other. What this does say is that the answer depends on the data, so it has to be measured on yours.
Unknown questions broke every threshold
CLINC150 also contains questions no intent covers, such as asking an assistant something outside its domain. A system with no "none of these" option has to hand those to a person, and the only signal it has is a low probability.
| CLINC150 test, thresholds for 95% accuracy | Answered automatically | Correct among those | Out-of-scope answered (of 1,000) | To a person |
|---|---|---|---|---|
| Classifier alone | 91.7% | 84.8% | 592 | 8.3% |
| Classifier, then Jev (top 5) | 96.8% | 81.7% | 844 | 3.2% |
| Classifier, then Jev (all 150 options) | 95.2% | 82.9% | 749 | 4.8% |
| Classifier, then Nimble (top 5) | 95.1% | 82.4% | 766 | 4.9% |
Every setting chosen for 95% came in around 82–85% on test. On the in-scope questions alone, the classifier's automatic answers were 96.1% correct; the whole shortfall is unknown questions answered as known ones. The cause is the validation set: 100 of its 3,100 messages are out of scope (3%), against 1,000 of 5,500 in the test set (18%). Thresholds tuned on a set with few unknowns do not expect many, and a decision model makes it worse, because it is good at finding the closest intent for any message, including one that has none. At the 98% target the pattern was the same and smaller: 207 leaks for the classifier alone, between 190 and 243 with a second stage.
With unknown questions in validation
CLINC150's training split also has 250 out-of-scope messages that the classifier never trains on. I added them to validation only, so thresholds are chosen knowing unknowns exist (10.4% of validation). Training, test and the classifier stay the same. This is a setting I added after seeing the result above, so it is shown beside the original rather than instead of it.
| CLINC150 test, classifier alone | Official validation (3% unknown) | With 250 more unknowns (10.4%) |
|---|---|---|
| 95% target: correct among automatic answers | 84.8% | 92.0% |
| 95% target: out-of-scope answered (of 1,000) | 592 | 271 |
| 95% target: to a person | 8.3% | 18.2% |
| 98% target: correct among automatic answers | 93.6% | 97.3% |
| 98% target: out-of-scope answered | 207 | 80 |
| 98% target: to a person | 20.8% | 30.3% |
Leaks fell by more than half, and the person's share rose by about ten points: that is the system handing unknown questions on, which is the job. It still fell short of the target, because 10.4% unknowns in validation is still fewer than 18% in test. If the share of unknown questions in your traffic changes, the thresholds have to be set again.
In this setting the second stage took work off the person in every Jev configuration at the 95% target, and the longer the list, the more: with all 150 options 1.9 points of the traffic (from 18.2% to 16.3%; paired bootstrap interval 1.2 to 2.6 points), with the top 10 1.2 points (0.8 to 1.7), with the top 5 0.8 points (0.2 to 1.4), and in none of them was the change in wrong answers distinguishable from zero (all 150: +0.05 points, −0.5 to +0.6). At the 98% target only Jev with all 150 options had any answers accepted, taking 2.1 points off (1.4 to 2.8) with 0.3 points more wrong answers (0.0 to 0.6). Nimble took 0.8 points off with the top 10 (0.04 to 1.5) and with the top 5, where the interval includes zero (−0.02 to 1.6); Tev1 took nothing off.
Shortlist or full list
Two measurements point in different directions here. Asked to label every message on its own, the decision models were slightly more accurate with a short list: on BANKING77 Nimble was right 85.4% of the time with the classifier's top 5, 84.4% with the top 10 and 83.2% with the top 20, and Jev 82.1% with the top 5 against 79.9% with all 77 options (92.8% against 92.1% on CLINC150). These are small differences and I did not test them. As the second stage of the system, though, the full list did better: with unknown questions in validation, Jev with all 150 options took the most work off the person (1.9 points, against 1.2 for the top 10 and 0.8 for the top 5), and with the official validation set it let fewer unknown questions through (749 against 844 for the top 5).
Speed
The classifier took 8 ms per message on four CPU threads. A second-stage call with five candidates took a median of 165–237 ms for the local models on an A100 that another job was also using (Nimble with 10 and 20 candidates took 281 and 385 ms on BANKING77), and Jev took 216–221 ms from Korea, with five candidates or the full list. Since the second stage only sees the messages the classifier hands on, 7–34% of traffic in these settings, most messages are answered in milliseconds.
What to do with this
- Start with the classifier and a person. Choose its threshold on held-out data. On both datasets it carried most of the load on its own.
- Put unknown questions in your validation set, in roughly the share you expect in traffic. Without them, the threshold that looks safe lets most unknown questions through.
- Add a decision model only if it beats the classifier on the messages the classifier hands on. Measure that on your own handed-on messages, same items, before you run it for everything. On BANKING77 it would have added latency and cost for nothing.
- Measure the list length as part of the system, not on its own. A shortlist made each model's own picks slightly more accurate, but the full list gave the most hand-offs saved and fewer unknown questions answered.
What this does not show
Two English datasets, one classifier, one version of each model, label names without descriptions. Thresholds for the second stage were searched on a grid in steps of 0.05; Jev's probabilities come rounded to two decimals, which limits how finely its threshold can be set. Some validation decisions rest on small samples (83 handed-on messages for BANKING77 at the 95% target, 33 in-scope ones for CLINC150). Latency for the local models was measured while another job shared the GPU. The validation set with added unknowns is a setting chosen after the first result, and it still holds fewer unknowns than the test set. None of this measures cost per message or what a person's time is worth; that is the next question.
The book's first steps are already readable: chapter 2, the first baseline, with its code.
I'm writing a book that builds this system step by step
One example from start to finish on BANKING77: a fixed split, a classifier, decision models, handing off what it doesn't know, and running it. Leave your email and I'll tell you once, when it's out.
Setup: classifier sentence-transformers/all-MiniLM-L6-v2 with scikit-learn logistic regression (C = 10), temperature-scaled on validation (T = 0.75 for BANKING77, 0.63 for CLINC150, fitted on in-scope validation messages by negative log-likelihood). BANKING77 split: 9,000 training and 1,003 validation messages from the 10,003 training set, 10% per intent, seed 20261001; test set 3,080. CLINC150 "plus": 15,000 in-scope training messages; official validation (3,000 + 100 out-of-scope) and test (4,500 + 1,000 out-of-scope); the variant adds the 250 out-of-scope messages of the training split to validation. Second stage through /v1/systemone: Nimble (nimble:latest) and Tev1 (tev1:4b) on Ollama 0.35.0 with OLLAMA_LLM_LIBRARY=cuda_v13 on one A100 80GB, Jev jev-1.13.0 through api.typesafe.ai, four requests at a time. Instructions as in the earlier posts; criteria are label names only. 94,614 second-stage calls in total, none failed. Thresholds: the classifier's on its own validation probabilities, the second stage's on a grid from 0 to 1 in steps of 0.05; the chosen setting maximises the share answered automatically on validation with automatic accuracy at or above the target. Confidence intervals for accuracy are Wilson; comparisons on the same handed-on messages are exact McNemar tests; differences between whole systems are paired bootstrap intervals over test messages (2,000 resamples). Scripts: scripts/jev-bench/route_cascade.py; results: `drafts/jev-bench/results/route-.json`. Measured 2026-10-02.*
Courses that go with this post
SOTAAZ courseCourses are sold one at a time. If you want all 13, there is a $199 lifetime bundle.
Subscribe to Newsletter
Related Posts

Jeff vs Jev on the Same Questions: Overall Scores and a 26-Option Limit (Fixed in v1.1)
On Jeff's own 4,599 questions, Jev scored 85.7 and Jeff-2B 83.0. Jeff v1.0 never picked an option past the 26th in my tests; v1.1 fixes that, remeasured.

Kev vs Jev: Kev 0.8B to 9B Benchmarked for Accuracy and Calibration
Kev (0.8B, 4B, 9B) and Jev 1.13 on the same 1,629 labelled messages (2,329 requests). On the three datasets Kev trained on, Kev-9B was as accurate or more (TREC 93.8% vs 89.0%). On three it never saw, Jev was ahead on two (CLINC150 68.5% vs 62.0%, MASSIVE 82.9% vs 76.0%). Jev's probabilities ran high and come rounded to two decimals, so on BANKING77 no threshold reached 95% accuracy.

How Many Labels Is a Decision Model Worth? Kev on Three Tasks It Never Saw
On CLINC150, MASSIVE and financial-news tweets, none of which Kev trained on, Kev-9B matched a small classifier trained on roughly 2 to 20 labelled examples per class. Its probabilities ran too low: stated confidence sat 10–17 points under its accuracy (calibration error 11–17%, against 2–9% on familiar data), so a 95%-accuracy threshold passed only 23–58% of messages. Out-of-scope questions got low probabilities: 5 of 100 passed that threshold. The 0.8B model answered 'Financials' to 152 of 200 tweets.
