Models & Algorithms••KR

Ollama's Decision Models on the Same Questions as Jev: Nimble and Tev1, Measured

On the same questions as Jev, Ollama's Nimble 9B scored 95.6% on TREC (Jev 89.0%) but 76.0 against 85.7 on a 4,599-question reasoning-heavy panel.

Ollama's Decision Models on the Same Questions as Jev: Nimble and Tev1, Measured

Ollama's Decision Models on the Same Questions as Jev: Nimble and Tev1, Measured

Ollama 0.35, released on 28 September, added a /v1/systemone endpoint that follows Jev's API: you send a state and typed questions, and a local model returns a probability for every option instead of text. It ships with three models: Nimble (9B, from Bespoke Labs) and Tev1 (4B and 0.8B, from Together AI), all fine-tuned from Qwen3.5.

I already had Jev, Kev and Jeff measured on the same messages, so I pointed the same harness at Ollama and ran all three models on them: five classification conditions from the earlier posts, and the 4,599-question panel from Jeff vs Jev.

The short answer

  • On short classification tasks, Nimble matched or beat Jev. TREC: 95.6% against 89.0% (p < 0.001). AG News and financial-news tweets: within three points either way, not significant.
  • On the reasoning-heavy panel, Jev was well ahead: 85.7 against Nimble's 76.0, with the largest gaps on WinoGrande and BBH.
  • Ollama's endpoint refused questions with more than 26 options for all three models, with a clear error rather than a wrong answer. BANKING77 (77 intents) cannot be asked as one question.
  • Tev1 0.8B trailed on TREC and on the panel overall (41% and 59.7), though it had the top AG News score with descriptions. Both Tev1 sizes reject inputs longer than about 2,050 tokens.
Heatmap of accuracy, same questions for every model. Rows: TREC names, TREC descriptions, AG News names, AG News descriptions, financial tweets, Jeff's panel. Jev 1.13: 89.0, 93.6, 89.0, 90.0, 71.0, 85.7. Nimble 9B: 95.6, 96.2, 90.0, 90.0, 68.0, 76.0. Tev1 4B: 71.2, 89.8, 88.5, 90.5, 67.5, 70.8. Tev1 0.8B: 41.0, 41.2, 86.5, 91.0, 46.0, 59.7. Kev-9B: 93.8, 96.6, 89.0, 90.5, 69.5, 73.8. Jeff-2B v1.1: 69.4, 87.0, 88.5, 88.5, 62.0, 81.9.

What each model saw in training

This decides how to read every number below, so it comes first.

Nimble publishes its training set: 2,676 synthetic examples across ten subject areas such as commerce, travel and software, with no named public datasets. I matched our test messages against it: none of the 500 TREC questions, 200 AG News articles or 200 financial tweets appear, and the dataset names do not occur in it. As far as its fine-tuning goes, Nimble's results here are zero-shot. What the Qwen3.5-9B base saw before that is not published.

Tev1 lists its sources, and two of them are ours: 1,500 AG News examples and 3,000 BANKING77 examples, alongside MultiNLI, BoolQ and SST-5. AG News is familiar ground for Tev1; TREC and the financial tweets are not.

Jev's training data is not published, so I cannot say either way for it.

Short classification: Nimble matched or beat Jev

Jev 1.13Nimble 9BTev1 4BTev1 0.8BKev-9BJeff-2B v1.1
TREC, label names (500)89.095.671.241.093.869.4
TREC, with descriptions (500)93.696.289.841.296.687.0
AG News, label names (200)89.090.088.586.589.088.5
AG News, with descriptions (200)90.090.090.591.090.588.5
Financial tweets, 20 topics (200)71.068.067.546.069.562.0

Against Jev, message by message, Nimble got 38 TREC questions right that Jev missed and missed 5 that Jev got (p < 0.001); with descriptions, 17 against 4 (p = 0.007). On AG News and the financial tweets the differences were within a few items and not significant (p ≥ 0.34). Kev-9B, which was trained on TREC, scored about the same as Nimble, which was not.

Tev1 4B gained 19 points on TREC when each label got a one-line description, the largest jump of any model here, and stayed level with the others on AG News, which it was trained on. Tev1 0.8B fell to 41% on TREC with or without descriptions.

The reasoning-heavy panel: Jev was well ahead

The panel is the one Jeff publishes: BBH, Financial PhraseBank, JudgeBench, RAGTruth and WinoGrande, 4,599 questions. None of its benchmarks is in Nimble's training set.

Jev 1.13Nimble 9BNimble vs Jev (only one right)
BBH (750)90.570.321 / 173, p < 0.001
Financial PhraseBank (999)85.384.558 / 66, p = 0.53
JudgeBench (350)79.464.321 / 74, p < 0.001
RAGTruth (1,500)81.878.9191 / 235, p = 0.037
WinoGrande (1,000)90.671.820 / 208, p < 0.001
All 4,59985.776.0311 / 756, p < 0.001

Nimble tied Jev on financial sentiment and trailed it everywhere reasoning matters. Nimble's panel total of 76.0 sits between Kev-9B (73.8) and Jeff-2B v1.1 (81.9). Among the Ollama models, Tev1 had the strongest financial sentiment results: on Financial PhraseBank Tev1 4B scored 93.9 and Tev1 0.8B 89.7, both above Jev's 85.3 (103 against 17, p < 0.001; 137 against 93, p = 0.004). SST-5, a sentiment dataset, is in Tev1's training data. Tev1 scored 70.8 (4B) and 59.7 (0.8B), but those totals count 138 rejected questions as wrong; more on that below.

Two limits you will hit

More than 26 options is refused. Sending BANKING77 with all 77 intents, Ollama 0.35's /v1/systemone returned HTTP 400: question "intent": criteria must contain 2–26 candidates for all three models, and CLINC150 and MASSIVE got the same: identical wording for every model, so it looks like a request check in the endpoint rather than in the models. That is the safer behaviour: Jeff v1.0 accepted such lists and silently never picked past the 26th (fixed in v1.1). But it means a 77-way intent task needs a shortlist step in your code before Ollama sees it. Tev1's own README asks for 2 to 24 options, since that is what it was trained on.

Tev1 rejects long inputs. On the panel, 138 questions came back with errors such as prompt 0 has 3429 tokens; expected 1–2050 (input is never truncated): 116 of JudgeBench's 350 and 22 of RAGTruth's 1,500. Refusing is better than truncating silently, but it caps what you can send. On the JudgeBench questions it did answer, Tev1 4B was right 63.2% of the time; Nimble, which has no such limit here, answered all of them.

Probabilities: well-behaved on TREC, over-confident on reasoning

Ollama returns full-precision probabilities, not the two-decimal rounding Jev's API returns, so thresholds can be set finely. How far the stated probability sat from actual accuracy depended on the task:

Average top probability minus accuracyJev 1.13Nimble 9BTev1 4B
TREC, label names−0.1+1.0+6.4
AG News, label names+6.7+4.9+1.6
Financial tweets+15.8+17.2+9.1
JudgeBench (panel)+0.4+21.1
WinoGrande (panel)+2.7+21.0

On TREC, Nimble's probabilities were almost exact, and a threshold could accept all 500 questions while keeping accepted answers 95% correct. On JudgeBench and WinoGrande it said it was about 21 points surer than it was. Nimble's card says no temperature was fitted after training, and Tev1's README says its "logprobs are model preferences, not calibrated confidence." Both are honest about it; the practical upshot is to set thresholds on your own labelled data, per task.

Speed

On one A100 with nothing else running, one request at a time over local HTTP, the median per TREC question was 192 ms for Nimble, 158 ms for Tev1 4B and 53 ms for Tev1 0.8B. On the same machine Kev-9B answered in 29 ms and Jeff in about 45 ms, so Ollama's path is slower here; I did not investigate why. Jev's API took about 220 ms from Korea; its server-side time was about 53 ms, measured on 30 BANKING77 requests on 29 September.

One setup note. On this server Ollama listed the A100 twice, once through Vulkan and once through CUDA, and chose Vulkan. Setting OLLAMA_LLM_LIBRARY=cuda_v13 put it on CUDA. Accuracy was identical either way; if your NVIDIA card feels slow under Ollama 0.35, check the server log for which backend it picked.

Which to use

  • Short local classification, no data leaving the machine: Nimble. On TREC it matched the best model here, on AG News and the financial tweets it was level with Jev, and its probabilities were well calibrated on TREC.
  • Financial sentiment: Tev1 4B scored highest of the Ollama models on Financial PhraseBank, 93.9 against Jev's 85.3.
  • Judging answers, multi-step reasoning, pronoun resolution: Jev, by 15 to 20 points on the same questions.
  • More than 26 options: shortlist first, whichever Ollama model you use.
  • Tev1 0.8B: only for tasks you have measured it on. It did well on AG News, which it trained on, and poorly on TREC, which it did not.

What this does not show

Five short-text conditions and one panel, all English. One version of each model (Ollama's Q8_0 builds) and of Jev (1.13). Nimble's TREC lead is zero-shot with respect to its fine-tuning data, but its base model's pretraining is unknown, and so is Jev's training data. Thresholds were searched on the same samples they describe. Latency is local HTTP on an A100 and will differ on a Mac or a consumer GPU.

Setup: Ollama 0.35.0 (Linux amd64 release tarball, CUDA 13 library, OLLAMA_LLM_LIBRARY=cuda_v13), /v1/systemone on 127.0.0.1, one A100 80GB PCIe. Models from the Ollama library: nimble:latest (blob bbf1d6fc03bb, 9.53 GB, Q8_0), tev1:4b (35f9281a3df5, 4.48 GB), tev1:0.8b (fa9732e3924d, 0.81 GB). Same request bodies, data files, label names, descriptions and instructions as the Kev vs Jev and Jeff vs Jev posts; Jev, Kev and Jeff rows are read from those result files. The panel is Jeff's jeff-panel output, 4,599 rows. Nimble's training set is bespokelabsai/nimble data/train.jsonl at 62076b4; overlap was checked by exact match of each test message, and an 80-character span of each panel question, after lowercasing and removing punctuation (scripts/jev-bench/check_nimble_overlap.py; 0 hits in all four sets). Rejected requests count as wrong in totals. Calibration gap is the mean top-option probability minus accuracy. McNemar tests are exact and two-sided. The accuracy runs shared the GPU with another job; latency was measured afterwards on an idle GPU, and both gave the same answers. Measured 2026-10-01.

Courses that go with this post

SOTAAZ course

Courses are sold one at a time. If you want all 13, there is a $199 lifetime bundle.

Stay Updated

Follow us for the latest posts and tutorials

Subscribe to Newsletter

Related Posts

How Many Labels Is a Decision Model Worth? Kev on Three Tasks It Never Saw
Models & Algorithms

How Many Labels Is a Decision Model Worth? Kev on Three Tasks It Never Saw

On CLINC150, MASSIVE and financial-news tweets, none of which Kev trained on, Kev-9B matched a small classifier trained on roughly 2 to 20 labelled examples per class. Its probabilities ran too low: stated confidence sat 10–17 points under its accuracy (calibration error 11–17%, against 2–9% on familiar data), so a 95%-accuracy threshold passed only 23–58% of messages. Out-of-scope questions got low probabilities: 5 of 100 passed that threshold. The 0.8B model answered 'Financials' to 152 of 200 tweets.

Kev vs MoJev: Two More Open Jev Alternatives, Accuracy and Confidence Measured
Models & Algorithms

Kev vs MoJev: Two More Open Jev Alternatives, Accuracy and Confidence Measured

Kev (0.8B to 9B) and MoJev (0.85B) both speak Jev's API. On BANKING77, TREC and AG News with human labels, Kev's probabilities were usable: at 95% accuracy it could auto-accept 53–61% of BANKING77, where laya managed 0%. But Kev was trained on these three datasets, and a small classifier trained on the same data still beat it on BANKING77. MoJev reports 0.79% calibration error; here it was 6–22%, with probabilities too low.

Can You Trust a Decision Model's Confidence? Repeated Answers and Probabilities, Measured
Models & Algorithms

Can You Trust a Decision Model's Confidence? Repeated Answers and Probabilities, Measured

Decision models sell a probability with every answer, so you can act on the sure ones and escalate the rest. I checked both ways of judging 'sure' on human-labelled data. Across 8 noise draws, 77–80% of openjev's wrong answers were unanimous. Probabilities did better, but how useful they were varied sharply by task: laya could auto-accept 88% of TREC at 95% accuracy and 0% of BANKING77. The new CLM-8B stayed near chance on all three datasets when given label names.