Classifier, LLM or Decision Model? A Measured Guide to Text Classification
One path through every text-classification measurement on this blog: on the same 154 banking messages, a CPU classifier scored 90.3%, an LLM with five retrieved examples 94.8%, and Jev 76.0%. Which to use, and when.

Classifier, LLM or Decision Model? A Measured Guide to Text Classification
Since mid-September I have measured a dozen ways to put a label on a short text: trained classifiers, frontier LLMs, local models, and the new decision models that return a probability for every option, such as Jev. Each post answered one question. This page puts them in the order you would meet those questions when building a system: what to try first, when a language model helps, which decision model to pick, and when to trust its probabilities.
If you want the history and architecture of text classifiers, from bag-of-words to transformers to Jev-style scoring heads, Sebastian Raschka's visual guide covers it well. This page is about choosing, with numbers measured on the same messages.
One table first
Most of the posts share one test: 154 messages from BANKING77, a public set of customer messages to a bank with 77 intents (two per intent, fixed seed). The same messages, the same 77 label names, scored against human labels:
| BANKING77, 154 messages | Correct | Labelled examples used | Post |
|---|---|---|---|
| GPT-5.6 Terra with five retrieved training examples | 94.8% (146) | 10,003 to search | Only retrieval moved it |
| MiniLM embeddings + logistic regression, on a CPU | 90.3% (139) | 10,003 | Label-only baselines |
| GPT-5.6 Terra, label names only | 83.8% (129) | 0 | Label-only baselines |
| Kev-9B (trained on BANKING77) | 82.5% (127) | in its training | Kev vs MoJev |
| Jev 1.13 | 76.0% (117) | 0 | Kev vs Jev |
| openjev (DiffusionGemma), one read | 66.9% | 0 | Jev alternatives |
| Jeff v1.1, 0.8B and 2B | 66.2% and 64.9% | 0 | Jeff vs Jev |
| Nimble 9B and Tev1 (Ollama 0.35) | refused: more than 26 options | 0 | Ollama's decision models |
The classifier scored 139 in the first post and 138 when rerun on a GPU to save per-message output; the GPT-5.6 Terra labels-only configuration scored between 125 and 129 across four runs. With 154 messages, a gap of a few points is not significant on its own; each post reports the paired test.
Two things in this table held up in the other posts. With many options and enough labels, a small classifier was hard to beat; with few options it was not (on TREC's six, Nimble scored 95.6% against the classifier's 90.4%). And when an LLM is given labelled examples, retrieving the nearest few helps more than anything else I tried.
Step 1: train the baseline first
If you have labelled examples, a sentence-embedding model with logistic regression on top is the first thing to try. It trains in seconds, answers in under 10 ms on a CPU, and on BANKING77 it beat every model that was not shown examples: two frontier LLMs, a grammar-constrained Qwen3-8B and three open Jev alternatives (label-only baselines, Jev alternatives).
Step 2: how many labels is a model worth?
The baseline needs labels; a decision model does not. On three tasks Kev had never seen, Kev-9B matched a classifier trained on roughly 2 to 20 examples per class. On the same tasks Jev was ahead of Kev-9B on two of three (CLINC150 68.5% against 62.0%, MASSIVE 82.9% against 76.0%; Kev vs Jev), which puts it about where a classifier trained on 20 per class sits:
| Same 400 and 175 messages | CLINC150 (max 75%) | MASSIVE |
|---|---|---|
| Classifier, 5 labels per class | 64.8% | 71.2% |
| Classifier, 20 labels per class | 69.8% | 82.3% |
| Classifier, all labels | 72.5% | 84.6% |
| Jev 1.13, none | 68.5% | 82.9% |
The 5- and 20-label rows are means over five draws of the examples, which varied by one to five points, so this is a placement, not a paired test. CLINC150's maximum is 75% because a quarter of the sample is out-of-scope questions that no option fits. If you can label a few dozen examples per class, the classifier is cheaper and as accurate; if you cannot, a decision model is the fastest start.
Step 3: if you use an LLM, give it examples, not rules
Holding GPT-5.6 Terra fixed on BANKING77, I changed only the prompt: shuffled labels, grouped labels, more or less reasoning. None of those moved accuracy significantly. Attaching the five nearest training examples fixed 19 messages and broke none (Only retrieval moved it). Rules written from the model's own mistakes did not do the same: 40 of them fixed 4 messages and broke 5 (Paper of the Week #4).
Step 4: choosing a decision model
All of these take the same request, a text plus questions with fixed options, and return a probability per option. They differ in what they were trained on and in how many options they accept.
- Jev, TypeSafe's hosted API, was ahead of Kev-9B on two of the three tasks Kev never saw and led every open model on reasoning-heavy questions: 85.7 on Jeff's 4,599-question panel (Jeff vs Jev). What it is and how it works: Jev in 10 minutes.
- Kev (0.8B to 9B, open weights) was as accurate as Jev or more on the datasets it trained on (TREC 93.8% against 89.0%) and behind on two of three it did not (Kev vs Jev, Kev vs MoJev).
- Nimble 9B on Ollama scored 95.6% on TREC against Jev's 89.0%, without TREC in its fine-tuning data, but 76.0 against 85.7 on the reasoning-heavy panel (Ollama's decision models).
- Jeff (0.8B and 2B) was better than Jev on financial sentiment and on checking answers against a source, and behind on reasoning (Jeff vs Jev).
- laya, openjev, NanoJev and DiffusionGemma were fast but well behind on 77 options (Jev alternatives, DiffusionGemma).
Check the option count before anything else. Several of these stop at 26: Ollama's endpoint refuses more, DiffusionGemma's example server caps at 26, and Jeff v1.0 accepted longer lists but never picked past the 26th (fixed in v1.1). Above that, you either shortlist candidates first or use a model that takes the full list.
Check the training data second. Kev was trained on BANKING77, TREC and AG News, Tev1 on AG News and BANKING77, laya on AG News. A model's score on a dataset it trained on says little about your task.
Step 5: when to act on the probability
The reason to use a decision model rather than a label is the probability: accept the confident answers, send the rest to a person. Whether that works depended more on the task than on the model.
- On BANKING77 Kev's probabilities let it auto-accept 53–61% of messages while keeping accepted answers 95% correct; laya could accept none (Kev vs MoJev).
- Jev's probabilities ran higher than its accuracy and come rounded to two decimals, so on BANKING77 no threshold kept accepted answers 95% correct (Kev vs Jev).
- On tasks Kev had not seen, its probabilities ran 10 to 17 points too low, and only 5 of 100 out-of-scope questions passed the 95% threshold, which is the behaviour you want for questions no option fits (Kev on unseen tasks).
- Asking the same model several times does not tell you when it is wrong: 77–80% of openjev's wrong answers were the same in all eight draws (Can you trust a decision model's confidence?).
Whatever the model, set the threshold on a few hundred of your own labelled messages, per task.
What these posts do not cover
All of it is English, short texts, and single-label classification. None of it measures a system running over time: how accuracy moves when the messages change, how to re-check thresholds, or what routing a share of traffic to a person costs. The posts measured one version of each model, and several used samples of 154 to 500 messages, so differences of a few points are only as reliable as the paired test reported with them.
Every number on this page comes from the linked post, where the setup, sample, seed and paired tests are given. BANKING77 sample: 154 test messages, two per intent, seed 20260918. CLINC150 sample: 300 in-scope test questions plus 100 out-of-scope. MASSIVE: 175 English test utterances.
Courses that go with this post
SOTAAZ courseCourses are sold one at a time. If you want all 13, there is a $199 lifetime bundle.
Subscribe to Newsletter
Related Posts

Kev vs MoJev: Two More Open Jev Alternatives, Accuracy and Confidence Measured
Kev (0.8B to 9B) and MoJev (0.85B) both speak Jev's API. On BANKING77, TREC and AG News with human labels, Kev's probabilities were usable: at 95% accuracy it could auto-accept 53–61% of BANKING77, where laya managed 0%. But Kev was trained on these three datasets, and a small classifier trained on the same data still beat it on BANKING77. MoJev reports 0.79% calibration error; here it was 6–22%, with probabilities too low.

Ollama's Decision Models on the Same Questions as Jev: Nimble and Tev1, Measured
On the same questions as Jev, Ollama's Nimble 9B scored 95.6% on TREC (Jev 89.0%) but 76.0 against 85.7 on a 4,599-question reasoning-heavy panel.

Jeff vs Jev on the Same Questions: Overall Scores and a 26-Option Limit (Fixed in v1.1)
On Jeff's own 4,599 questions, Jev scored 85.7 and Jeff-2B 83.0. Jeff v1.0 never picked an option past the 26th in my tests; v1.1 fixes that, remeasured.
