Models & Algorithms••KR

Which Messages Should Go to a Person? A Classifier, a Decision Model and a Hand-Off, Measured

On BANKING77 a decision model after a classifier saved no hand-offs. On CLINC150 unknown questions broke the thresholds; adding 250 to validation halved the leaks.

Which Messages Should Go to a Person? A Classifier, a Decision Model and a Hand-Off, Measured

Which Messages Should Go to a Person? A Classifier, a Decision Model and a Hand-Off, Measured

The last step of the text-classification guide is acting on a probability: answer the confident cases automatically and send the rest to a person. This post builds that system and measures it end to end. It has three stages:

  1. A small classifier (MiniLM embeddings with logistic regression, on a CPU; the same baseline as before) answers when its probability is above a threshold.
  2. Otherwise the classifier's top five candidates go to a decision model, which answers when its own probability is above a second threshold. I tried Nimble 9B and Tev1 4B locally through Ollama 0.35 (measured on their own here), and Jev 1.13 through its API, also with the full list of options.
  3. Otherwise a person. A message handed to a person is never counted as answered correctly.

The question is what the second stage buys: how many messages it takes off the person's desk at the same accuracy.

How the thresholds were chosen

Every setting was chosen on a validation set and scored once on a test set the selection never saw. For BANKING77, 77 intents of messages to a bank, I held out 10% of the training set (1,003 messages, stratified by intent) for validation and scored on the full test set of 3,080. For CLINC150, 150 intents for a virtual assistant, I used the official splits: validation 3,100 messages, test 5,500, and both contain out-of-scope questions that fit no intent. The thresholds are the ones that answered the most validation messages while keeping automatic answers at 95% (or 98%) correct. The test result is then reported as it came out, including when it falls short of the target.

The short answer

  • On BANKING77, no decision model reduced the person's share. The messages the classifier was unsure about were just as hard for Nimble, Tev1 and Jev. At the 95% target the classifier alone answered 93.2% of the test set automatically, 96.2% of those correctly, and every cascade ended up with the same numbers or no better.
  • On CLINC150, Jev and Nimble did help with the uncertain messages: at the 98% target Jev got 78.8% of them right where the classifier's first choice got 61.6% (p < 0.001); Tev1 was not clearly better.
  • Unknown questions broke every threshold. CLINC150's validation set is 3% out-of-scope; its test set is 18%. Thresholds chosen for 95% accuracy produced 84.8% on test, with 592 of 1,000 out-of-scope questions answered as if they were in scope. Adding a decision model let more through.
  • Putting unknown questions into validation cut the leaks by more than half. With 250 more out-of-scope messages in validation, leaks fell from 592 to 271, and Jev as the second stage then took 1.9 points of the traffic off the person without adding errors. The test still fell short of the 95% target.
Slope chart, CLINC150 test set, out-of-scope questions answered automatically out of 1,000, thresholds chosen for 95% accuracy. With the official validation set (3% out-of-scope) and with 250 out-of-scope messages added to validation (10.4%): classifier alone 592 to 271, classifier then Jev with all 150 options 749 to 269, classifier then Nimble 766 to 284, classifier then Tev1 757 to 274.

BANKING77: the second stage had nothing to add

The classifier on its own was already strong: 92.9% of test messages right, and the correct intent was among its top five candidates for 99.3% of them. At the 95% target it answered 2,871 of 3,080 test messages automatically, 96.2% correctly (95% confidence interval 95.4–96.8%), and handed 209 to a person. At the 98% target it answered 82.9% at 98.0% (97.4–98.5%) and handed off 17.1%.

The second stage only sees the messages the classifier hands on, so the fair comparison is on those same messages: the classifier's first choice against each model's pick.

BANKING77 test, messages the classifier handed on95% target (209)98% target (526)
Classifier's first choice47.4%67.9%
Nimble 9B, top 552.2% (p = 0.31)60.3% (p = 0.003)
Tev1 4B, top 551.2% (p = 0.46)54.9% (p < 0.001)
Jev 1.13, top 551.2% (p = 0.43)57.0% (p < 0.001)
Jev 1.13, all 77 options47.8% (p = 1.0)55.5% (p < 0.001)

At the 95% target the models were no better than the classifier's own guess, and at the 98% target they were worse. Their probabilities did not separate right from wrong well enough either: on the validation messages the classifier handed on, Nimble's answers with a probability of 0.99 or more were right 58–82% of the time depending on the setting, out of only 10 to 30 such answers. So the threshold search on validation chose to accept no second-stage answers at all, for every model except Jev with all 77 options, which accepted answers at a probability of exactly 1.00. On test that added 35 automatic answers, a difference in automatic share of +0.3 points (−0.2 to +0.7) that is indistinguishable from none.

Two caveats sit with this table. Tev1 lists 3,000 BANKING77 examples in its training data, and even so it did not beat the classifier here. And the "accept nothing" choice was made on 83 validation messages at the 95% target and 195 at 98%, so it is a decision on a small sample.

CLINC150: the second stage helped on the uncertain messages

The same comparison on CLINC150 came out the other way.

CLINC150 test, in-scope messages the classifier handed on95% target (48)98% target (349)
Classifier's first choice39.6%61.6%
Nimble 9B, top 566.7% (p = 0.011)69.6% (p = 0.018)
Tev1 4B, top 558.3% (p = 0.049)64.5% (p = 0.45)
Jev 1.13, top 568.8% (p = 0.007)78.8% (p < 0.001)

So "the messages a classifier is unsure of are hard for everyone" held on BANKING77 and not on CLINC150. These runs do not say why. The obvious guess, that BANKING77's uncertain cases are near-duplicate intents, does not separate the two: in both datasets, when the classifier was wrong on a handed-on message, the right intent was usually its second choice (57–60% of the time on BANKING77, 63% on CLINC150 at the 98% target), and the most common mix-ups look alike, such as verify_my_identity against why_verify_identity in one and user_name against what_is_your_name in the other. What this does say is that the answer depends on the data, so it has to be measured on yours.

Unknown questions broke every threshold

CLINC150 also contains questions no intent covers, such as asking an assistant something outside its domain. A system with no "none of these" option has to hand those to a person, and the only signal it has is a low probability.

CLINC150 test, thresholds for 95% accuracyAnswered automaticallyCorrect among thoseOut-of-scope answered (of 1,000)To a person
Classifier alone91.7%84.8%5928.3%
Classifier, then Jev (top 5)96.8%81.7%8443.2%
Classifier, then Jev (all 150 options)95.2%82.9%7494.8%
Classifier, then Nimble (top 5)95.1%82.4%7664.9%

Every setting chosen for 95% came in around 82–85% on test. On the in-scope questions alone, the classifier's automatic answers were 96.1% correct; the whole shortfall is unknown questions answered as known ones. The cause is the validation set: 100 of its 3,100 messages are out of scope (3%), against 1,000 of 5,500 in the test set (18%). Thresholds tuned on a set with few unknowns do not expect many, and a decision model makes it worse, because it is good at finding the closest intent for any message, including one that has none. At the 98% target the pattern was the same and smaller: 207 leaks for the classifier alone, between 190 and 243 with a second stage.

With unknown questions in validation

CLINC150's training split also has 250 out-of-scope messages that the classifier never trains on. I added them to validation only, so thresholds are chosen knowing unknowns exist (10.4% of validation). Training, test and the classifier stay the same. This is a setting I added after seeing the result above, so it is shown beside the original rather than instead of it.

CLINC150 test, classifier aloneOfficial validation (3% unknown)With 250 more unknowns (10.4%)
95% target: correct among automatic answers84.8%92.0%
95% target: out-of-scope answered (of 1,000)592271
95% target: to a person8.3%18.2%
98% target: correct among automatic answers93.6%97.3%
98% target: out-of-scope answered20780
98% target: to a person20.8%30.3%

Leaks fell by more than half, and the person's share rose by about ten points: that is the system handing unknown questions on, which is the job. It still fell short of the target, because 10.4% unknowns in validation is still fewer than 18% in test. If the share of unknown questions in your traffic changes, the thresholds have to be set again.

In this setting the second stage took work off the person in every Jev configuration at the 95% target, and the longer the list, the more: with all 150 options 1.9 points of the traffic (from 18.2% to 16.3%; paired bootstrap interval 1.2 to 2.6 points), with the top 10 1.2 points (0.8 to 1.7), with the top 5 0.8 points (0.2 to 1.4), and in none of them was the change in wrong answers distinguishable from zero (all 150: +0.05 points, −0.5 to +0.6). At the 98% target only Jev with all 150 options had any answers accepted, taking 2.1 points off (1.4 to 2.8) with 0.3 points more wrong answers (0.0 to 0.6). Nimble took 0.8 points off with the top 10 (0.04 to 1.5) and with the top 5, where the interval includes zero (−0.02 to 1.6); Tev1 took nothing off.

Shortlist or full list

Two measurements point in different directions here. Asked to label every message on its own, the decision models were slightly more accurate with a short list: on BANKING77 Nimble was right 85.4% of the time with the classifier's top 5, 84.4% with the top 10 and 83.2% with the top 20, and Jev 82.1% with the top 5 against 79.9% with all 77 options (92.8% against 92.1% on CLINC150). These are small differences and I did not test them. As the second stage of the system, though, the full list did better: with unknown questions in validation, Jev with all 150 options took the most work off the person (1.9 points, against 1.2 for the top 10 and 0.8 for the top 5), and with the official validation set it let fewer unknown questions through (749 against 844 for the top 5).

Speed

The classifier took 8 ms per message on four CPU threads. A second-stage call with five candidates took a median of 165–237 ms for the local models on an A100 that another job was also using (Nimble with 10 and 20 candidates took 281 and 385 ms on BANKING77), and Jev took 216–221 ms from Korea, with five candidates or the full list. Since the second stage only sees the messages the classifier hands on, 7–34% of traffic in these settings, most messages are answered in milliseconds.

What to do with this

  • Start with the classifier and a person. Choose its threshold on held-out data. On both datasets it carried most of the load on its own.
  • Put unknown questions in your validation set, in roughly the share you expect in traffic. Without them, the threshold that looks safe lets most unknown questions through.
  • Add a decision model only if it beats the classifier on the messages the classifier hands on. Measure that on your own handed-on messages, same items, before you run it for everything. On BANKING77 it would have added latency and cost for nothing.
  • Measure the list length as part of the system, not on its own. A shortlist made each model's own picks slightly more accurate, but the full list gave the most hand-offs saved and fewer unknown questions answered.

What this does not show

Two English datasets, one classifier, one version of each model, label names without descriptions. Thresholds for the second stage were searched on a grid in steps of 0.05; Jev's probabilities come rounded to two decimals, which limits how finely its threshold can be set. Some validation decisions rest on small samples (83 handed-on messages for BANKING77 at the 95% target, 33 in-scope ones for CLINC150). Latency for the local models was measured while another job shared the GPU. The validation set with added unknowns is a setting chosen after the first result, and it still holds fewer unknowns than the test set. None of this measures cost per message or what a person's time is worth; that is the next question.

The book's first steps are already readable: chapter 2, the first baseline, with its code.

I'm writing a book that builds this system step by step

One example from start to finish on BANKING77: a fixed split, a classifier, decision models, handing off what it doesn't know, and running it. Leave your email and I'll tell you once, when it's out.

One email at launch, nothing else.

Setup: classifier sentence-transformers/all-MiniLM-L6-v2 with scikit-learn logistic regression (C = 10), temperature-scaled on validation (T = 0.75 for BANKING77, 0.63 for CLINC150, fitted on in-scope validation messages by negative log-likelihood). BANKING77 split: 9,000 training and 1,003 validation messages from the 10,003 training set, 10% per intent, seed 20261001; test set 3,080. CLINC150 "plus": 15,000 in-scope training messages; official validation (3,000 + 100 out-of-scope) and test (4,500 + 1,000 out-of-scope); the variant adds the 250 out-of-scope messages of the training split to validation. Second stage through /v1/systemone: Nimble (nimble:latest) and Tev1 (tev1:4b) on Ollama 0.35.0 with OLLAMA_LLM_LIBRARY=cuda_v13 on one A100 80GB, Jev jev-1.13.0 through api.typesafe.ai, four requests at a time. Instructions as in the earlier posts; criteria are label names only. 94,614 second-stage calls in total, none failed. Thresholds: the classifier's on its own validation probabilities, the second stage's on a grid from 0 to 1 in steps of 0.05; the chosen setting maximises the share answered automatically on validation with automatic accuracy at or above the target. Confidence intervals for accuracy are Wilson; comparisons on the same handed-on messages are exact McNemar tests; differences between whole systems are paired bootstrap intervals over test messages (2,000 resamples). Scripts: scripts/jev-bench/route_cascade.py; results: `drafts/jev-bench/results/route-.json`. Measured 2026-10-02.*

Courses that go with this post

SOTAAZ course

Courses are sold one at a time. If you want all 13, there is a $199 lifetime bundle.

Stay Updated

Follow us for the latest posts and tutorials

Subscribe to Newsletter

Related Posts