What Are You Deciding, and Which Mistakes Are Expensive? Chapter 1 of a Book on Building Decision Systems
Free chapter: before training anything, count what each intent can tell you, find the intents likely to be confused, and write down what each mistake costs. The code runs in under a second.

What Are You Deciding, and Which Mistakes Are Expensive? Chapter 1 of a Book on Building Decision Systems
This is the first chapter of a book I am writing, Building a Decision System with Small Models. It follows one example from start to finish: sorting messages to a bank into 77 intents, answering the ones the system is sure of and handing the rest to a person. Chapter 2 is already out as a sample. The code for both chapters is an 11 KB download: decision-book-code-ch01.zip. The table of contents is at the end.
Before training anything, this chapter writes down three things the rest of the book scores against:
- what the system decides: one of 77 intents, or "hand this to a person";
- which intents it is likely to mix up, found from the training messages alone;
- what each kind of mistake costs, as a small table in
cost.pythat every later chapter imports.
No model is trained here and nothing needs a GPU. Once the data is downloaded, the chapter's script runs in under a second (0.8 s on my machine).
1.1 Two decisions, not one
A bank receives "I am still waiting on my card?" and has to route it. The obvious decision is which of the 77 intents it is: card_arrival, here. The second decision is easy to miss: whether to answer at all. A system that may hand a message to a person can be wrong in two ways that cost different amounts:
- it answers, and the answer is wrong;
- it hands the message on when it could have answered it.
A classifier's accuracy treats every wrong intent the same and has no notion of handing on. So before comparing any models, we decide how to count. This book counts in hand-offs:
| What happened to the message | Cost |
|---|---|
| answered automatically and right | 0 |
| handed to a person | 1 |
| answered automatically and wrong | r × severity |
r is how many hand-offs one wrong answer is worth. severity makes some wrong answers worse than others (1.4). You will not know r exactly, and you do not need to: the book scores across a range of r and asks where the decision changes.
1.2 Count before you measure
Set up the environment as in the README (the same one serves every chapter), then run:
python ch01_define.pyThe first lines are counts:
train 9000 val 1003 test 3080 intents 77
messages per intent
train 31 to 168
val 4 to 19
test 40 to 40The split is the one made in bank.py and used in every chapter (chapter 2 explains why it is fixed before anything is trained). What matters here is the middle line: some intents have only four validation messages. Every choice in this book is made on validation, so it is worth knowing what four, or forty, messages can tell you. The script prints the 95% interval around an accuracy of 90% at a few sample sizes:
| Messages | 95% interval around 90% | Width |
|---|---|---|
| 5 | 46% to 99% | 53 points |
| 13 | 64% to 98% | 34 points |
| 40 | 77% to 96% | 19 points |
| 100 | 83% to 94% | 12 points |
| 400 | 87% to 93% | 6 points |
(These are Wilson intervals, the ones used for proportions throughout the book.) So a per-intent accuracy on the test set, with its 40 messages per intent, is good to within about ten points either way, and on validation it is mostly noise. Comparisons in this book are therefore made on whole sets, or on groups of intents, not intent by intent.
1.3 Which intents will be confused?
Some intents are close. card_arrival ("Where is my card?") and card_delivery_estimate ("How long does a card take to arrive?") are different labels for messages that often read the same. You can find candidates before you have a model: describe each intent by its average training message and see which averages are close.
The script does this with TF-IDF, word-and-character counts similar to chapter 2's first baseline (which uses characters 3 to 5 long; here 2 to 5):
features = make_union(TfidfVectorizer(ngram_range=(1, 2), sublinear_tf=True),
TfidfVectorizer(analyzer="char_wb", ngram_range=(2, 5), sublinear_tf=True))
X = features.fit_transform([r["text"] for r in train])
centroids = np.vstack([np.asarray(X[y == lab].mean(axis=0)) for lab in labels])
centroids /= np.linalg.norm(centroids, axis=1, keepdims=True)
sim = centroids @ centroids.T # cosine similarity between every pair of intentsOnly the training messages are used. The five most similar pairs:
| Similarity | Intents |
|---|---|
| 0.854 | verify_my_identity / why_verify_identity |
| 0.828 | card_payment_wrong_exchange_rate / wrong_exchange_rate_for_cash_withdrawal |
| 0.782 | disposable_card_limits / get_disposable_virtual_card |
| 0.760 | unable_to_verify_identity / why_verify_identity |
| 0.760 | getting_virtual_card / virtual_card_not_working |
Read the list as questions to ask, not answers. Some pairs are near-duplicates that people would also disagree on; some differ in one word that matters ("not working"). When I checked such a list against a trained classifier's actual mistakes on the test set (Which mistakes are expensive?), the 77 most similar pairs covered about half of its 219 errors, and the costliest pair, a payment the customer does not recognise against a compromised card, ranked 177th of 2,926. A list made from text alone is a start. Chapter 2 gives you the real mix-ups.
1.4 Which mistakes are expensive?
Not all wrong answers are equal. Open cost.py. It holds two judgements that the code cannot make for you:
Intents where the customer is at risk right now. Their money or their account: a lost or stolen card, a compromised card, a card payment, cash withdrawal or direct debit they did not make, a charge taken twice, a card the machine swallowed, the wrong amount of cash, a transfer that never arrived, a transfer they want to cancel. Eleven intents. A wrong answer to one of these costs five times as much.
Teams. Each intent belongs to one of eight teams that would handle it: card delivery, card problems, card payments, cash, top-ups, transfers, exchange, account. A wrong answer that lands with the right team costs a fifth as much, because a colleague passes it along instead of the customer starting again.
From these come four scenarios, which later chapters report side by side:
| Scenario | What a wrong answer costs |
|---|---|
uniform | r, always |
at-risk | 5r if the message is at risk, otherwise r |
same-team | 0.2r if the answer is in the right team, otherwise r |
both | at-risk first, then same-team |
The numbers 5 and 0.2 are mine. Yours will differ, and so will your list of at-risk intents and your teams. What matters is when you write them down: before you score any system. Once you have seen results, it is easy to pick a cost table that makes the system you like come first. Writing it first, and scoring across a range rather than at one value, removes that temptation.
The script checks the table against the data before it does anything else (cost.check(labels)): every intent must have a team, and every at-risk intent must exist. It also prints each at-risk intent's closest neighbour:
cancel_transfer -> transfer_not_received_by_recipient 0.535 same team
card_swallowed -> declined_cash_withdrawal 0.527 OTHER TEAM
compromised_card -> lost_or_stolen_card 0.533 same team
...Ten of the eleven have their nearest neighbour inside the same team. The one that does not, card_swallowed, is a pair to watch.
1.5 Where a decision flips
Here is the cost table in use. In chapter 5 you will build a system that answers when it is sure, hands the rest on, and needs a threshold. I measured a system of that kind on this blog, trained on all 77 intents (Which messages should go to a person?); on the 3,080 test messages, two of its thresholds gave:
| Setting | Wrong answers | Handed to a person |
|---|---|---|
| answers more | 109 | 209 |
| answers less | 50 | 526 |
Which is better? With the uniform cost, the total for each is wrong × r + handed, so they cost the same when
109r + 209 = 50r + 526, that is, r = 317 ÷ 59 = 5.37.
cost.break_even() does this arithmetic, and the script prints both costs at a few values of r:
r = 2.00: 0.139 vs 0.203 hand-offs per message -> answers more
r = 5.00: 0.245 vs 0.252 hand-offs per message -> answers more
r = 5.37: 0.258 vs 0.258 hand-offs per message -> same
r = 10.00: 0.422 vs 0.333 hand-offs per message -> answers lessIf a wrong answer costs you more than about five hand-offs, the stricter setting is cheaper. The first setting was in fact chosen for "95% of automatic answers correct", and that is what an accuracy target hides: it picks an r for you without saying which.
You do not need to know your r to use this. You need to know whether it is above or below 5.37.
1.6 A checklist for the evaluation set
The script ends by checking the data against four rules:
[x] every intent has at least 40 test messages
[ ] every intent has at least 20 validation messages
[ ] the test set holds questions outside the intents
[x] the at-risk intents have at least 100 test messages togetherTwo fail, and both matter later. Validation is thin per intent, so thresholds in chapter 5 are chosen on the whole set, and a threshold for very expensive mistakes rests on very few errors. And BANKING77 has no messages outside its 77 intents, while real traffic does. Chapter 5 deals with that by holding ten intents out of training, so that their messages play the questions the system has never seen.
For your own data, the questions behind the four lines are:
- How many messages does each option have, in the set you choose on and in the set you report on?
- Are the expensive options represented well enough to measure?
- Does the evaluation set contain messages that fit no option, in roughly the share you expect?
- Did you write the cost table before you scored anything?
Exercises
- Your own r. Think of one wrong answer your system could give that would be expensive and one that would be cheap. Roughly how many hand-offs is each worth? Is 5.37 inside that range?
- Change the at-risk list. Add
pin_blockedandpasscode_forgottentoAT_RISKincost.py(a customer locked out of their account). Doescard_swallowedstill have a neighbour in another team? Which at-risk intents now do? - A different similarity. Replace the TF-IDF features with character n-grams only (
analyzer="char_wb",ngram_range=(2, 5)). Which of the top ten pairs stay?
When it does not run
cost table does not match the labels. You changedcost.pyand an intent is missing fromTEAMS, or a name is misspelt. The message lists which.unexpected contents (sha256 …). The data file is incomplete or was changed. Delete thedata/folder and run again.ModuleNotFoundError: sklearn. The environment is not active. Runsource .venv/bin/activate(Windows:.venv\Scripts\activate) in the book's folder first.
Measured with the code in this chapter: Python 3.11.4, scikit-learn 1.9.1, numpy 2.4.6. Every number in this chapter is in measured/ch01.json, except the two settings in 1.5 and the coverage of mix-ups in 1.3, which come from the posts linked there. Data: BANKING77 at commit 57ec275; split seed 20261001.
The rest of the book
| Chapter | The question it answers | What you have at the end |
|---|---|---|
| 1. What are you deciding? | What are the options, and which mistakes are expensive? | Label definitions, an evaluation set, a cost for each kind of mistake |
| 2. The first baseline | How far does a simple classifier get? | A fixed train/validation/test split and two CPU classifiers at about nine in ten |
| 3. Using a language model (API) | How do you give an LLM examples and descriptions? | An LLM classifier with retrieved examples, compared on the same messages |
| 4. Decision models (API, or local with a GPU) | When is a dedicated decision model worth it? | A same-message comparison report against your classifier |
| 5. Handing off what it doesn't know | Which answers do you accept automatically, and which go to a person? | A routing system with thresholds chosen on validation and tested once. BANKING77 has no questions outside its 77 intents, so ten intents are held out of training and their messages play the unknown questions |
| 6. Running it | What changes when the messages and the intents change? | The ten held-out intents arrive as new ones: what the system does with them, what to re-check after labelling them, and the logs that tell you when |
Chapters 1, 2, 5 and 6 run on one machine without a GPU; chapters 3 and 4 call hosted APIs (an API key and a few dollars), with a local alternative through Ollama for readers who have a GPU.
Get told when the book is out
The other four chapters are being written in the same way: one example, code you can run, every number from it. Leave your email and I'll tell you once, when it's out.
Courses that go with this post
SOTAAZ courseCourses are sold one at a time. If you want all 13, there is a $199 lifetime bundle.
Subscribe to Newsletter
Related Posts

Which Messages Should Go to a Person? A Classifier, a Decision Model and a Hand-Off, Measured
On BANKING77 a decision model after a classifier saved no hand-offs. On CLINC150 unknown questions broke the thresholds; adding 250 to validation halved the leaks.

Classifier, LLM or Decision Model? A Measured Guide to Text Classification
One path through every text-classification measurement on this blog: on the same 154 banking messages, a CPU classifier scored 90.3%, an LLM with five retrieved examples 94.8%, and Jev 76.0%. Which to use, and when.

Jev-Style Decision Models in Ollama 0.35: Nimble and Tev1, Setup and Benchmark
Jev itself is not in Ollama, but Ollama 0.35 adds a Jev-compatible API with Nimble and Tev1. How to call /v1/systemone, its 26-option limit, and how the models scored against Jev on the same questions.
