How Far Does a Simple Classifier Get? Chapter 2 of a Book on Building Decision Systems
Free sample chapter: split BANKING77 before training, then two CPU classifiers reach 91.2% and 92.9% on the test set, with code that runs in about a minute.

How Far Does a Simple Classifier Get? Chapter 2 of a Book on Building Decision Systems
This is the second chapter of a book I am writing, Building a Decision System with Small Models. It follows one example from start to finish: sorting messages to a bank into 77 intents, answering the ones the system is sure of and handing the rest to a person. The chapter stands on its own, and its code is a 5 KB download: decision-book-code-ch02.zip. The table of contents is at the end.
By the end of this chapter you will have three things that the rest of the book builds on:
- a fixed split of the data into training, validation and test sets, made before anything is trained;
- a classifier that sorts customer messages into 77 intents, runs on a laptop CPU and answers in milliseconds;
- a list of the mistakes it makes, which is where the later chapters start.
Everything here runs without a GPU. On my machine the chapter's script, including the first downloads, ran in about a minute after a one-minute install; a laptop will take a few minutes.
2.1 The task
A bank receives short messages such as "I am still waiting on my card?" and has to decide what each one is about before anyone answers it. BANKING77, a public dataset from PolyAI, collects 13,083 such messages, each labelled by a person with one of 77 intents: card_arrival, exchange_rate, verify_my_identity and so on. It is published under CC BY 4.0, so you can use it in your own work with attribution.
The dataset comes in two files: 10,003 training messages and 3,080 test messages. The training file is not balanced. After the split below, the rarest intent has 31 training messages and the most common 168.
Why this dataset for a whole book? It looks like a real routing problem. There are many options, several of them close to each other ("a card payment I don't recognise" against "a direct debit I don't recognise"), and the messages are short and informal. It is also small enough that every experiment in the book runs on one machine in minutes.
2.2 Setting up
You need Python 3.11 or newer. In a new folder with the book's code:
python3 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install --upgrade pip
pip install torch==2.8.0 --index-url https://download.pytorch.org/whl/cpu
pip install -r requirements.txtThe third command installs PyTorch's CPU build, which is a fraction of the size of the default one. The versions are pinned so that your numbers come out close to the ones in this chapter. On my machine the install took under a minute.
2.3 Split before you train
Open bank.py. Its load() function downloads the two files the first time (and checks them against a fixed hash, so you know you have the same data as the book), then splits them:
- train: 90% of each intent's training messages (9,000 messages). Models learn from these.
- validation: the other 10% of each intent (1,003 messages). Every choice we make, which model, which setting, and from chapter 5 on which threshold, is made by looking at these.
- test: the 3,080 official test messages. We score them once, at the end, with the choices already made.
Taking 10% of each intent keeps every intent in the validation set, including the rare ones. The random seed is fixed, so the split is the same every time you run the code and in every chapter.
Why bother with three sets now, when we have not trained anything? Because the habit is easiest to keep from the start. If you try ten settings and keep the one that does best on the test set, the test score is no longer an honest estimate: you chose it. Later in the book we choose thresholds, and that mistake becomes expensive. When I measured a routing system this way on another dataset, CLINC150 (Which messages should go to a person?), thresholds chosen on validation fell well short of their target on a test set with more unknown questions; without a separate test set scored once, that would have gone unnoticed.
2.4 Baseline one: word counts
Run:
python ch02_baseline.pyThe first model is the oldest trick in text classification. TF-IDF turns each message into a long vector of weighted word and character counts; logistic regression learns one weight per feature per intent. In the code it is a few lines:
tfidf = make_pipeline(
make_union(TfidfVectorizer(ngram_range=(1, 2), sublinear_tf=True),
TfidfVectorizer(analyzer="char_wb", ngram_range=(3, 5), sublinear_tf=True)),
LogisticRegression(C=10, max_iter=3000),
)
tfidf.fit(train_texts, train_labels)The word features catch phrases such as "exchange rate"; the character features (3 to 5 letters) catch spelling variants and typos such as "recieved". On validation it was right 89.6% of the time, and it answered one message in 1.5 ms. It needs no download and trained in 22 seconds.
2.5 Baseline two: sentence embeddings
The second model replaces word counts with a sentence embedding: a small pretrained network (all-MiniLM-L6-v2, 22 million parameters, about 90 MB) turns each message into 384 numbers that place messages with similar meaning near each other. "Where is my new card?" and "My card hasn't arrived yet" share almost no words but end up close together. Logistic regression then works on those 384 numbers.
encoder = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2", device="cpu")
E_train = encoder.encode(train_texts, normalize_embeddings=True)
emb = LogisticRegression(C=10, max_iter=3000).fit(E_train, train_labels)On validation it was right 90.8% of the time and took 6.6 ms per message, most of it running the encoder. Training took 9 seconds, nearly all of it encoding the 9,000 training messages.
| On validation (1,003 messages) | Accuracy | Right intent in the top 5 | One message (CPU) |
|---|---|---|---|
| TF-IDF + logistic regression | 89.6% | 98.4% | 1.5 ms |
| Embeddings + logistic regression | 90.8% | 99.0% | 6.6 ms |
The two are close. On 1,003 messages a gap of 1.2 points is 12 messages, so do not read much into which one is ahead; chapter 4 shows how to compare two models on the same messages properly. What matters more is where both already are: about nine in ten, with no language model.
2.6 Reading the mistakes
The script prints the most common mix-ups. For the embedding model on validation:
3 x transfer_not_received_by_recipient -> predicted transfer_timing
3 x direct_debit_payment_not_recognised -> predicted card_payment_not_recognised
2 x pending_cash_withdrawal -> predicted declined_cash_withdrawal
2 x contactless_not_working -> predicted card_arrival
2 x unable_to_verify_identity -> predicted verify_my_identityMost of these are neighbours: two intents a person would also have to think about. Some are arguably labelling choices rather than model errors. One validation message labelled direct_debit_payment_not_recognised reads "There is a payment in my app that I did not make. I have not used that card all day…"; the model said card_payment_not_recognised, and a reader could too. The other column of the table says the same thing from another side: the right intent was among the model's top five guesses for 99.0% of validation messages. When the model is wrong, it is usually wrong by a little.
That observation does not tell you what to do yet, and it would be easy to jump to a conclusion. Chapter 5 tests one tempting idea: handing the messages the classifier is unsure of, with its top candidates, to a decision model. When I tried it on this blog, it helped on CLINC150 and not on BANKING77, so the answer depends on the data.
2.7 The test set, once
Now that the choices are made (we will carry both models forward), score the test set:
python ch02_baseline.py --test| On test (3,080 messages) | Accuracy | Right intent in the top 5 |
|---|---|---|
| TF-IDF + logistic regression | 91.2% | 99.0% |
| Embeddings + logistic regression | 92.9% | 99.3% |
Both scored higher on test than on validation. I have not found out why. The validation messages come from the training file and the test messages from a separate file, so differences between the two files are where to look; it is not a sign that something leaked, since the test set played no part in any choice. It is a first reminder that a score belongs to the set it was measured on.
2.8 Where this stands
On this blog I compared other approaches on a sample of 154 of these test messages, two per intent (text-classification guide). On that same sample, this embedding classifier (trained there on all 10,003 training messages) got 90.3% right, GPT-5.6 Terra given only the 77 label names between 81.2% and 83.8% over four runs, and Jev, a hosted decision model, 76.0%. With 154 messages each of those is a few messages either way. The direction is the point: a classifier trained on the labelled messages you already have is a strong first answer, and it costs nothing per message.
What it cannot do is tell you which of its answers to trust. Every message gets an answer, including the ones it is unsure of and the ones no intent fits. That is the subject of the rest of the book.
Exercises
- Change the regularisation. In
ch02_baseline.py, tryC=1andC=100for both models. Compare on validation only. Does the ranking of the two models change? - Fewer labels. Train the embedding model on 20 messages per intent instead of all of them (sample them from
trainwith a fixed seed). How much accuracy do you lose? Chapter 4 compares this with decision models that need no labelled examples at all. - Read ten mistakes. Print ten validation messages the embedding model got wrong, with the true and predicted intents. For each, decide whether the model is wrong, the label is wrong, or both intents fit. Keep the list; chapter 5 comes back to messages like these.
When it does not run
ensurepip is not availableor an old Python. Some Linux systems ship Python withoutvenv. Install Python 3.11 or newer from python.org, or with your package manager (python3.11-venvon Debian and Ubuntu).- PyTorch downloads several gigabytes. You skipped the CPU install line, and pip fetched the GPU build. Remove the environment and start again with the commands in 2.2.
- The model will not download (no internet, or a proxy). Run once on a connected machine, then copy the Hugging Face cache folder (
~/.cache/huggingface) across, or setHF_HOMEto a folder you can copy. unexpected contents (sha256 …). The data file is incomplete or was changed. Delete thedata/folder and run again.- Your numbers differ slightly from the book's. Check the package versions with
pip listagainstrequirements.txt. Small differences in the last digit can come from a different CPU; differences of a point or more mean something else changed.
Measured with the code in this chapter: Python 3.11.4, torch 2.8.0 (CPU build), scikit-learn 1.9.1, sentence-transformers 6.1.0, numpy 2.4.6, on an AMD EPYC 7742 (64 cores) with the GPU hidden, starting from an empty model cache. Timings are medians over 200 validation messages answered one at a time, with PyTorch's default of 64 CPU threads (on 4 threads the same classifier took 8 ms in the routing measurements on this blog), and will be slower on a laptop. Every number in this chapter is in measured/ch02.json. Data: BANKING77 at commit 57ec275; split seed 20261001.
The rest of the book
| Chapter | The question it answers | What you have at the end |
|---|---|---|
| 1. What are you deciding? | What are the options, and which mistakes are expensive? | Label definitions, an evaluation set, a cost for each kind of mistake |
| 2. The first baseline | How far does a simple classifier get? | A fixed train/validation/test split and two CPU classifiers at about nine in ten |
| 3. Using a language model (API) | How do you give an LLM examples and descriptions? | An LLM classifier with retrieved examples, compared on the same messages |
| 4. Decision models (API, or local with a GPU) | When is a dedicated decision model worth it? | A same-message comparison report against your classifier |
| 5. Handing off what it doesn't know | Which answers do you accept automatically, and which go to a person? | A routing system with thresholds chosen on validation and tested once. BANKING77 has no questions outside its 77 intents, so ten intents are held out of training and their messages play the unknown questions |
| 6. Running it | What changes when the messages and the intents change? | The ten held-out intents arrive as new ones: what the system does with them, what to re-check after labelling them, and the logs that tell you when |
Chapters 1, 2, 5 and 6 run on one machine without a GPU; chapters 3 and 4 call hosted APIs (an API key and a few dollars), with a local alternative through Ollama for readers who have a GPU.
Get told when the book is out
The other five chapters are being written in the same way: one example, code you can run, every number from it. Leave your email and I'll tell you once, when it's out.
Courses that go with this post
SOTAAZ courseCourses are sold one at a time. If you want all 13, there is a $199 lifetime bundle.
Subscribe to Newsletter
Related Posts

Ollama's Decision Models on the Same Questions as Jev: Nimble and Tev1, Measured
On the same questions as Jev, Ollama's Nimble 9B scored 95.6% on TREC (Jev 89.0%) but 76.0 against 85.7 on a 4,599-question reasoning-heavy panel.

Jeff vs Jev on the Same Questions: Overall Scores and a 26-Option Limit (Fixed in v1.1)
On Jeff's own 4,599 questions, Jev scored 85.7 and Jeff-2B 83.0. Jeff v1.0 never picked an option past the 26th in my tests; v1.1 fixes that, remeasured.

Kev vs Jev: Kev 0.8B to 9B Benchmarked for Accuracy and Calibration
Kev (0.8B, 4B, 9B) and Jev 1.13 on the same 1,629 labelled messages (2,329 requests). On the three datasets Kev trained on, Kev-9B was as accurate or more (TREC 93.8% vs 89.0%). On three it never saw, Jev was ahead on two (CLINC150 68.5% vs 62.0%, MASSIVE 82.9% vs 76.0%). Jev's probabilities ran high and come rounded to two decimals, so on BANKING77 no threshold reached 95% accuracy.
