Models & Algorithms•SOTAAZ Lab••KR

We Trained Our Own Decision Model with Unsloth: 75% on BANKING77, 56% When BANKING77 Is Left Out of Training

We followed Unsloth's recipe for turning Qwen3.5-0.8B into a decision model (4.15 GB peak memory, about 20 minutes per run on a shared GPU). BANKING77 and CLINC150 matched Unsloth's table (75.4% and 75.0%), but those test rows come from datasets in the training mix. With the two datasets removed from training, the same 500 rows scored 56.5% and 55.9%. typed-decisions came out 11.8 points below the table; the recipe's decontamination step drops 850 of its 1,200 training cases, and in a post hoc run that kept them it reached 74.3%.

We Trained Our Own Decision Model with Unsloth: 75% on BANKING77, 56% When BANKING77 Is Left Out of Training

We Trained Our Own Decision Model with Unsloth: 75% on BANKING77, 56% When BANKING77 Is Left Out of Training

A decision model reads an input and a few questions about it, and instead of writing text it picks one of the options you gave it and says how sure it is. "Which team should handle this ticket: billing, technical or sales?" comes back as billing with a probability. We have measured several of these models on this blog, mostly ones other people trained.

This week Unsloth published a guide, "Train your own Decision Model with Unsloth," that turns an ordinary language model into one; the training code was merged into Unsloth on 7 October. Its table says Qwen3.5-0.8B goes from 7% to 74% on BANKING77 (77 banking intents) and from 19% to 76% on CLINC150 (150 intents plus out-of-scope), using about 4 GB of GPU memory. The numbers are Unsloth's own measurement.

That is an appealing claim for anyone who wants their own small classifier, so we tried it ourselves. Reading the guide next to the code, we also noticed something the table doesn't say: BANKING77 and CLINC150 are both in the training mix. The guide does say the test sets were "decontaminated," and we wanted to know what that covers.

We fixed three questions, the decision rule and the sentence for each outcome before training anything:

  1. Reproduce. With Unsloth's own script and defaults, do we get the table's numbers?
  2. Training scope. If BANKING77 and CLINC150 are removed from training, how much do the same 500 test rows change?
  3. Use it. Save the model and ask it a few decisions we wrote ourselves. This part is a demonstration, not a measurement.

How we trained it

The guide shows two settings: LoRA rank 64 for one epoch next to the table, and "epochs 2, rank 16" in a code example that trains a 4B model on one dataset. The repository has a script, scripts/train_decision_from_lm.py, whose defaults match the first: rank 64, one epoch, 12,000 mixture rows from the same 12 sources the guide lists, and the same test sets (2,000 typed-decisions decisions, 500 rows each from BANKING77 and CLINC150). We used that script unmodified, at Unsloth commit d632c82b.

  • Model: unsloth/Qwen3.5-0.8B, 16-bit weights (the script's default, no 4-bit loading).
  • Training data: 1,000 rows from each of 12 public datasets, plus typed-decisions, a benchmark of workflow decisions, for 755 steps of 16 rows.
  • Hardware: one A100 80GB shared with another training job. On that shared GPU each run took 17 to 21 minutes; the sharing makes these rough, so we don't compare times between runs. Peak memory allocated by PyTorch was 4.15 GB, in line with the guide's 4 GB.
  • Seeds: 3407 (the script's default), 1 and 2 for every condition. The test rows are the same in every run.
  • Rule: a difference counts only if it exceeds twice its standard error across the three seeds. For reproduction, our average had to be within 3 points of the table.

If you want to try it, the whole recipe is one command after installing Unsloth:

bash
python scripts/train_decision_from_lm.py --model unsloth/Qwen3.5-0.8B \
  --sources banking77,clinc150,mnli,snli,wanli,boolq,ag_news,sst5,mmlu,commonsense_qa,arc,prompt_injections \
  --out my-decision-model

Listing the 12 sources matters. The script's default list also has a 13th, a gated dataset that is skipped with a one-line message unless you have accepted its terms on Hugging Face, so the same command can train on different data on different machines.

Dot plot of accuracy for three test sets. For each set: before training (grey ring), Unsloth's recipe (blue, 3 seeds), the same recipe without BANKING77 and CLINC150 (orange, 3 seeds), and Unsloth's reported value (open diamond). typed-decisions: before 36.5%, recipe 61.2%, without the two datasets 59.8%, reported 73%. BANKING77: before 4.5%, recipe 75.4%, without 56.5%, reported 74%. CLINC150: before 19.7%, recipe 75.0%, without 55.9%, reported 76%.

1. Two of the table's three numbers came out the same

Test setBefore trainingAfter (our average, 3 seeds)Unsloth's table
typed-decisions (2,000 decisions)36.5%61.2%73%
BANKING77 (500 rows)4.5%75.4%74%
CLINC150 (500 rows)19.7%75.0%76%

Before training, our numbers are close to the table's starting points (36%, 7%, 19%). After training, BANKING77 and CLINC150 land within 1.4 points of the table. typed-decisions is 11.8 points lower, outside our 3-point margin, so the reproduction holds for two sets out of three.

We looked for the cause in the data first. The typed-decisions data files have not changed since 16 September, three weeks before the guide, so it isn't a newer version of the dataset. The cause turned out to be in the recipe itself; section 3 has it.

The table's fourth column, "Holdout acc 78%," came out at 82.1% for us. That number is measured on 10% of the training mix held back from training, so it includes BANKING77 and CLINC150 rows too.

2. Remove BANKING77 and CLINC150 from training

The guide says the test sets were decontaminated against the training data. In the code that means: drop any training row that shares a run of 13 consecutive words with a test row. BANKING77 and CLINC150 messages are short (median 10 and 9 words; 332 and 430 of the 500 test messages are under 13 words), so in practice the check removes only word-for-word copies. Of the 2,000 candidate rows the script draws from each dataset, it removed 3 from BANKING77 and none from CLINC150. The 1,000 rows each dataset contributes to training stay in, with the same label set as the test rows.

That is a fair way to report in-domain accuracy. It just measures something different from a task the model has never seen. So we ran the same recipe a second time with those two datasets removed from the source list, keeping the total at 12,000 rows (the other ten sources got 1,200 rows each instead of 1,000).

Test setRecipeWithout the two datasetsDifferenceTwice the SE
BANKING7775.4%56.5%−18.9 points2.3
CLINC15075.0%55.9%−19.1 points3.5
typed-decisions61.2%59.8%−1.4 points2.9

Both intent sets dropped by about 19 points, far more than seed changes move them. typed-decisions, which both conditions trained on, didn't change by more than seeds do.

So the table's BANKING77 and CLINC150 numbers come from training on the same datasets. Left out of training, the recipe scored about 56% on both. That 56% is still a real result: before training the model was at 4.5% on BANKING77, where picking at random among 77 intents gives 1.3%. The other ten sources taught it a good part of intent classification. If your task is intent classification over intents that weren't in training, 56% is the closer estimate.

3. Why typed-decisions came out lower

The same 13-word check runs on typed-decisions, and there it behaves differently. typed-decisions inputs are structured JSON records generated from templates, so a training case and a test case often share long runs of field names and boilerplate. We counted what the check removes from the 1,200 training cases:

typed-decisions workflowTraining cases dropped
Agent trace review300 of 300
Invoice processing300 of 300
Security incidents210 of 300
Customer service40 of 300

No test input is an exact copy of a training input. But after the check, the model trained on no agent-trace or invoice cases at all, and those two workflows are half of the typed-decisions test. In the agent-trace workflow, every test task sentence also appears in the training split; the dataset reuses its tasks across splits by design.

That is a likely reason for the 11.8-point gap, but we found it after seeing the result, so we tested it separately and label it post hoc. We reran the recipe with one change: typed-decisions training cases are checked only against the BANKING77 and CLINC150 test rows, so all 1,200 stay in. This follows the dataset's own train/test split, the same one its card uses for models trained on it.

Test setRecipeAll 1,200 typed-decisions cases kept (post hoc)DifferenceTwice the SEUnsloth's table
typed-decisions61.2%74.3%+13.1 points2.873%
BANKING7775.4%75.7%+0.3 points2.274%
CLINC15075.0%73.0%−2.0 points5.676%

Keeping the 850 cases raised typed-decisions by 13.1 points, to 74.3%, within 1.3 points of Unsloth's table. BANKING77 and CLINC150 didn't change by more than seeds move them. CLINC150's average happens to sit exactly 3 points below the table, on the edge of our margin.

Landing within 1.3 points of the table once these cases are kept fits the explanation that Unsloth's table came from a run that had them in training. We changed only this one thing and can't see Unsloth's run, so we don't claim that is what happened. What we can say: with the recipe exactly as published today, typed-decisions comes out near 61%; keep those cases and it comes out near the table.

4. Using the model we made

We saved one more run of the recipe and asked it five decisions we wrote before running it. These are single examples, not a measurement.

InputQuestionAnswer
"I was charged twice for invoice #4411. Please refund the duplicate today."Which team: billing, technical, sales?billing, 99.4%
The same message in KoreanSame questionbilling, 99.9%
A blog comment sharing a measurement, with a link to a GPU deals siteIs it promotional spam?yes, 55.7%
A deploy log: 3 of 40 health checks failing, error rate 0.1% to 2.4%Severity: no impact / can wait / today / page now"needs attention today," 79.6% (page now 19.4%)
"I ordered a new card two weeks ago and it still hasn't shown up."Our own three intentscard delivery, 99.2%

The Korean ticket is one example from a model trained on English data, so it shows only that this sentence worked. The spam comment mixes a real measurement with an ad, and the answer came out close to even, at 55.7%.

One more thing we noticed while saving it: the saved run used the same seed and arguments as one of our measured runs, and still came out 3.3 points lower on typed-decisions and 3.4 points higher on BANKING77. GPU training is not bit-for-bit repeatable here, so a single run's number can move by a few points. Train more than once before trusting a small difference.

What to check when you train your own

  • Is your test set's dataset in the training mix? The script's mix includes BANKING77 and CLINC150, and "decontaminated" only removes near-copies. Evaluate on data from a source you left out.
  • How many training rows survived decontamination? The script prints the row count. For typed-decisions it kept 350 of 1,200 cases, which was easy to miss.
  • Did a gated dataset get skipped? List your sources explicitly so every machine trains on the same data.
  • Run more than one seed. Our three seeds on BANKING77 spanned 73.4% to 76.8%.

Limits

  • Only Qwen3.5-0.8B, only Unsloth's script defaults. Larger models or more epochs could move every number here.
  • Three seeds per condition; differences under about 3 points are not visible.
  • In the run without BANKING77 and CLINC150, the other sources got 1,200 rows each instead of 1,000. That could account for a small part of the change, but not a 19-point drop that matches across both sets.
  • 500 test rows per intent dataset, the same rows in every run.
  • The typed-decisions explanation is post hoc, and the run that tests it changed one thing; we did not test other possible causes of the gap.
  • The demonstration answers are single examples.

The plan, scripts, per-run results and logs are in the reproduction package below. If we train a larger model or test the label-reading behavior from our label bias post on this one, the newsletter will have it.

Files for this post

Reproduction package

The harness, raw logs and result CSVs behind this post. No account needed; the README has a command that recomputes the tables without a GPU.

train-decision-model-unsloth-repro.zip · 172 KB

Download