Models & Algorithms••KR

How Far Does a Simple Classifier Get? Chapter 2 of a Book on Building Decision Systems

Free sample chapter: split BANKING77 before training, then two CPU classifiers reach 91.2% and 92.9% on the test set, with code that runs in about a minute.

How Far Does a Simple Classifier Get? Chapter 2 of a Book on Building Decision Systems

How Far Does a Simple Classifier Get? Chapter 2 of a Book on Building Decision Systems

This is the second chapter of a book I am writing, Building a Decision System with Small Models. It follows one example from start to finish: sorting messages to a bank into 77 intents, answering the ones the system is sure of and handing the rest to a person. The chapter stands on its own, and its code is a 5 KB download: decision-book-code-ch02.zip. The table of contents is at the end.

By the end of this chapter you will have three things that the rest of the book builds on:

  • a fixed split of the data into training, validation and test sets, made before anything is trained;
  • a classifier that sorts customer messages into 77 intents, runs on a laptop CPU and answers in milliseconds;
  • a list of the mistakes it makes, which is where the later chapters start.

Everything here runs without a GPU. On my machine the chapter's script, including the first downloads, ran in about a minute after a one-minute install; a laptop will take a few minutes.

2.1 The task

A bank receives short messages such as "I am still waiting on my card?" and has to decide what each one is about before anyone answers it. BANKING77, a public dataset from PolyAI, collects 13,083 such messages, each labelled by a person with one of 77 intents: card_arrival, exchange_rate, verify_my_identity and so on. It is published under CC BY 4.0, so you can use it in your own work with attribution.

The dataset comes in two files: 10,003 training messages and 3,080 test messages. The training file is not balanced. After the split below, the rarest intent has 31 training messages and the most common 168.

Why this dataset for a whole book? It looks like a real routing problem. There are many options, several of them close to each other ("a card payment I don't recognise" against "a direct debit I don't recognise"), and the messages are short and informal. It is also small enough that every experiment in the book runs on one machine in minutes.

2.2 Setting up

You need Python 3.11 or newer. In a new folder with the book's code:

bash
python3 -m venv .venv
source .venv/bin/activate            # Windows: .venv\Scripts\activate
pip install --upgrade pip
pip install torch==2.8.0 --index-url https://download.pytorch.org/whl/cpu
pip install -r requirements.txt

The third command installs PyTorch's CPU build, which is a fraction of the size of the default one. The versions are pinned so that your numbers come out close to the ones in this chapter. On my machine the install took under a minute.

2.3 Split before you train

Open bank.py. Its load() function downloads the two files the first time (and checks them against a fixed hash, so you know you have the same data as the book), then splits them:

  • train: 90% of each intent's training messages (9,000 messages). Models learn from these.
  • validation: the other 10% of each intent (1,003 messages). Every choice we make, which model, which setting, and from chapter 5 on which threshold, is made by looking at these.
  • test: the 3,080 official test messages. We score them once, at the end, with the choices already made.

Taking 10% of each intent keeps every intent in the validation set, including the rare ones. The random seed is fixed, so the split is the same every time you run the code and in every chapter.

Why bother with three sets now, when we have not trained anything? Because the habit is easiest to keep from the start. If you try ten settings and keep the one that does best on the test set, the test score is no longer an honest estimate: you chose it. Later in the book we choose thresholds, and that mistake becomes expensive. When I measured a routing system this way on another dataset, CLINC150 (Which messages should go to a person?), thresholds chosen on validation fell well short of their target on a test set with more unknown questions; without a separate test set scored once, that would have gone unnoticed.

2.4 Baseline one: word counts

Run:

bash
python ch02_baseline.py

The first model is the oldest trick in text classification. TF-IDF turns each message into a long vector of weighted word and character counts; logistic regression learns one weight per feature per intent. In the code it is a few lines:

python
tfidf = make_pipeline(
    make_union(TfidfVectorizer(ngram_range=(1, 2), sublinear_tf=True),
               TfidfVectorizer(analyzer="char_wb", ngram_range=(3, 5), sublinear_tf=True)),
    LogisticRegression(C=10, max_iter=3000),
)
tfidf.fit(train_texts, train_labels)

The word features catch phrases such as "exchange rate"; the character features (3 to 5 letters) catch spelling variants and typos such as "recieved". On validation it was right 89.6% of the time, and it answered one message in 1.5 ms. It needs no download and trained in 22 seconds.

2.5 Baseline two: sentence embeddings

The second model replaces word counts with a sentence embedding: a small pretrained network (all-MiniLM-L6-v2, 22 million parameters, about 90 MB) turns each message into 384 numbers that place messages with similar meaning near each other. "Where is my new card?" and "My card hasn't arrived yet" share almost no words but end up close together. Logistic regression then works on those 384 numbers.

python
encoder = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2", device="cpu")
E_train = encoder.encode(train_texts, normalize_embeddings=True)
emb = LogisticRegression(C=10, max_iter=3000).fit(E_train, train_labels)

On validation it was right 90.8% of the time and took 6.6 ms per message, most of it running the encoder. Training took 9 seconds, nearly all of it encoding the 9,000 training messages.

On validation (1,003 messages)AccuracyRight intent in the top 5One message (CPU)
TF-IDF + logistic regression89.6%98.4%1.5 ms
Embeddings + logistic regression90.8%99.0%6.6 ms

The two are close. On 1,003 messages a gap of 1.2 points is 12 messages, so do not read much into which one is ahead; chapter 4 shows how to compare two models on the same messages properly. What matters more is where both already are: about nine in ten, with no language model.

2.6 Reading the mistakes

The script prints the most common mix-ups. For the embedding model on validation:

3 x  transfer_not_received_by_recipient  ->  predicted transfer_timing
3 x  direct_debit_payment_not_recognised  ->  predicted card_payment_not_recognised
2 x  pending_cash_withdrawal  ->  predicted declined_cash_withdrawal
2 x  contactless_not_working  ->  predicted card_arrival
2 x  unable_to_verify_identity  ->  predicted verify_my_identity

Most of these are neighbours: two intents a person would also have to think about. Some are arguably labelling choices rather than model errors. One validation message labelled direct_debit_payment_not_recognised reads "There is a payment in my app that I did not make. I have not used that card all day…"; the model said card_payment_not_recognised, and a reader could too. The other column of the table says the same thing from another side: the right intent was among the model's top five guesses for 99.0% of validation messages. When the model is wrong, it is usually wrong by a little.

That observation does not tell you what to do yet, and it would be easy to jump to a conclusion. Chapter 5 tests one tempting idea: handing the messages the classifier is unsure of, with its top candidates, to a decision model. When I tried it on this blog, it helped on CLINC150 and not on BANKING77, so the answer depends on the data.

2.7 The test set, once

Now that the choices are made (we will carry both models forward), score the test set:

bash
python ch02_baseline.py --test
On test (3,080 messages)AccuracyRight intent in the top 5
TF-IDF + logistic regression91.2%99.0%
Embeddings + logistic regression92.9%99.3%

Both scored higher on test than on validation. I have not found out why. The validation messages come from the training file and the test messages from a separate file, so differences between the two files are where to look; it is not a sign that something leaked, since the test set played no part in any choice. It is a first reminder that a score belongs to the set it was measured on.

2.8 Where this stands

On this blog I compared other approaches on a sample of 154 of these test messages, two per intent (text-classification guide). On that same sample, this embedding classifier (trained there on all 10,003 training messages) got 90.3% right, GPT-5.6 Terra given only the 77 label names between 81.2% and 83.8% over four runs, and Jev, a hosted decision model, 76.0%. With 154 messages each of those is a few messages either way. The direction is the point: a classifier trained on the labelled messages you already have is a strong first answer, and it costs nothing per message.

What it cannot do is tell you which of its answers to trust. Every message gets an answer, including the ones it is unsure of and the ones no intent fits. That is the subject of the rest of the book.

Exercises

  1. Change the regularisation. In ch02_baseline.py, try C=1 and C=100 for both models. Compare on validation only. Does the ranking of the two models change?
  2. Fewer labels. Train the embedding model on 20 messages per intent instead of all of them (sample them from train with a fixed seed). How much accuracy do you lose? Chapter 4 compares this with decision models that need no labelled examples at all.
  3. Read ten mistakes. Print ten validation messages the embedding model got wrong, with the true and predicted intents. For each, decide whether the model is wrong, the label is wrong, or both intents fit. Keep the list; chapter 5 comes back to messages like these.

When it does not run

  • ensurepip is not available or an old Python. Some Linux systems ship Python without venv. Install Python 3.11 or newer from python.org, or with your package manager (python3.11-venv on Debian and Ubuntu).
  • PyTorch downloads several gigabytes. You skipped the CPU install line, and pip fetched the GPU build. Remove the environment and start again with the commands in 2.2.
  • The model will not download (no internet, or a proxy). Run once on a connected machine, then copy the Hugging Face cache folder (~/.cache/huggingface) across, or set HF_HOME to a folder you can copy.
  • unexpected contents (sha256 …). The data file is incomplete or was changed. Delete the data/ folder and run again.
  • Your numbers differ slightly from the book's. Check the package versions with pip list against requirements.txt. Small differences in the last digit can come from a different CPU; differences of a point or more mean something else changed.

Measured with the code in this chapter: Python 3.11.4, torch 2.8.0 (CPU build), scikit-learn 1.9.1, sentence-transformers 6.1.0, numpy 2.4.6, on an AMD EPYC 7742 (64 cores) with the GPU hidden, starting from an empty model cache. Timings are medians over 200 validation messages answered one at a time, with PyTorch's default of 64 CPU threads (on 4 threads the same classifier took 8 ms in the routing measurements on this blog), and will be slower on a laptop. Every number in this chapter is in measured/ch02.json. Data: BANKING77 at commit 57ec275; split seed 20261001.

The rest of the book

ChapterThe question it answersWhat you have at the end
1. What are you deciding?What are the options, and which mistakes are expensive?Label definitions, an evaluation set, a cost for each kind of mistake
2. The first baselineHow far does a simple classifier get?A fixed train/validation/test split and two CPU classifiers at about nine in ten
3. Using a language model (API)How do you give an LLM examples and descriptions?An LLM classifier with retrieved examples, compared on the same messages
4. Decision models (API, or local with a GPU)When is a dedicated decision model worth it?A same-message comparison report against your classifier
5. Handing off what it doesn't knowWhich answers do you accept automatically, and which go to a person?A routing system with thresholds chosen on validation and tested once. BANKING77 has no questions outside its 77 intents, so ten intents are held out of training and their messages play the unknown questions
6. Running itWhat changes when the messages and the intents change?The ten held-out intents arrive as new ones: what the system does with them, what to re-check after labelling them, and the logs that tell you when

Chapters 1, 2, 5 and 6 run on one machine without a GPU; chapters 3 and 4 call hosted APIs (an API key and a few dollars), with a local alternative through Ollama for readers who have a GPU.

Get told when the book is out

The other five chapters are being written in the same way: one example, code you can run, every number from it. Leave your email and I'll tell you once, when it's out.

The most-picked use case becomes the example in the next measured post.

One email at launch, nothing else.

Courses that go with this post

SOTAAZ course

Courses are sold one at a time. If you want all 13, there is a $199 lifetime bundle.

Stay Updated

Follow us for the latest posts and tutorials

Subscribe to Newsletter

Related Posts