Models & Algorithms••KR

Kev vs Jev: Kev 0.8B to 9B Benchmarked for Accuracy and Calibration

Kev (0.8B, 4B, 9B) and Jev 1.13 on the same 1,629 labelled messages (2,329 requests). On the three datasets Kev trained on, Kev-9B was as accurate or more (TREC 93.8% vs 89.0%). On three it never saw, Jev was ahead on two (CLINC150 68.5% vs 62.0%, MASSIVE 82.9% vs 76.0%). Jev's probabilities ran high and come rounded to two decimals, so on BANKING77 no threshold reached 95% accuracy.

Kev vs Jev: Kev 0.8B to 9B Benchmarked for Accuracy and Calibration

Kev vs Jev: Kev 0.8B to 9B Benchmarked for Accuracy and Calibration

Jev is TypeSafe's hosted decision model: you send text and a question with a fixed set of answers, and it returns one answer with a probability for each option. Kev is an open-weights model that accepts the same request format and runs on your own GPU.

In the last two posts I measured Kev on three datasets it trained on and three it never saw. This time I ran Jev 1.13 (jev-1.13.0) on the same rows: same messages, same label names, same one-line descriptions, same instructions, scored against the same human labels. Every comparison below is paired, message by message.

The short answer

  • On datasets Kev trained on, Kev-9B was as accurate as Jev or more. It was ahead on TREC and tied on AG News. On BANKING77 it got ten more right, but that gap is not statistically clear.
  • On datasets Kev never saw, Jev was more accurate on two of three. CLINC150: 68.5% against 62.0%. MASSIVE: 82.9% against 76.0%. On financial-news tweets the two were tied.
  • Jev's probabilities ran high, Kev's ran low on new tasks. Jev also returns them rounded to two decimals, and 43 of the 1,021 answers it gave at 1.00 were wrong. On BANKING77 that left no threshold that kept accepted answers 95% correct.
Dumbbell chart, Kev-9B in green and Jev 1.13 in red, on the same rows. Accuracy, tasks Kev trained on: BANKING77 82 vs 76, TREC names 94 vs 89, TREC with descriptions 97 vs 94, AG News 89 vs 89, AG News with descriptions 90 vs 90. Tasks Kev never saw: CLINC150 62 vs 68, MASSIVE 76 vs 83, financial tweets 70 vs 71. Share auto-accepted at 95% accuracy, Kev vs Jev: BANKING77 61 vs 0, TREC 98 vs 68, TREC with descriptions 100 vs 97, AG News 84 vs 76, AG News with descriptions 90 vs 0, CLINC150 43 vs 57, MASSIVE 58 vs 75, financial tweets 23 vs 0.

Accuracy on datasets Kev trained on

Kev's training data includes the train splits of BANKING77, TREC and AG News. My samples come from the test splits, so Kev never saw these exact messages, but these are the tasks it knows best. TypeSafe does not publish what Jev was trained on, so I cannot say whether Jev saw these public datasets.

Kev-9BJev 1.13Only one got it right (Kev / Jev)
BANKING77, 77 intents (154)127 (82.5%)117 (76.0%)17 / 7, p = 0.06
TREC, label names (500)469 (93.8%)445 (89.0%)33 / 9, p < 0.001
TREC, with descriptions (500)483 (96.6%)468 (93.6%)18 / 3, p = 0.001
AG News, label names (200)178 (89.0%)178 (89.0%)7 / 7, p = 1
AG News, with descriptions (200)181 (90.5%)180 (90.0%)6 / 5, p = 1

The p-values are exact McNemar tests on the messages only one of the two got right. Jev's 76.0% on BANKING77 is close to the 76.3% that nibzard's benchmark reported with all 77 intents on a different sample.

The small classifier from the first post, MiniLM embeddings with logistic regression trained on each dataset's train split, still beat both on BANKING77 (138 of 154; against Jev, 24 to 3, p < 0.001). On TREC with descriptions Jev beat it (468 against 452, 36 to 20, p = 0.04), and on AG News all three were within four articles.

Accuracy on datasets Kev never saw

These are the three tasks from the last post, chosen because they appear nowhere in Kev's training data. They are the fairer test of what you would get on your own task.

Kev-9BJev 1.13Only one got it right (Kev / Jev)
CLINC150, 150 intents + out-of-scope (400)248 (62.0%)274 (68.5%)5 / 31, p < 0.001
MASSIVE, 60 intents (175)133 (76.0%)145 (82.9%)5 / 17, p = 0.02
Financial tweets, 20 topics (200)139 (69.5%)142 (71.0%)11 / 14, p = 0.69

CLINC150's 100 out-of-scope questions have no correct option, so 75% is the ceiling there. On the 300 in-scope questions Jev was 91.3% accurate and Kev-9B 82.7%.

Last time, Kev-9B on these tasks matched a classifier trained on roughly 2 to 20 labelled examples per class. Jev's accuracy sits near the top of that range. The classifier trained on 20 examples per class averaged 69.8%, 82.3% and 66.9% over five draws, against Jev's 68.5%, 82.9% and 71.0%. I did not run paired tests against the classifier draws here, so read that as "about the same", not a ranking.

Every Kev size against Jev

Kev comes in 0.8B, 4B and 9B sizes that run on the same server, so the question is also which size is worth it. Accuracy on the same rows:

Kev-0.8BKev-4BKev-9BJev 1.13
BANKING7773.4%81.2%82.5%76.0%
TREC, names94.6%90.0%93.8%89.0%
TREC, descriptions96.4%95.4%96.6%93.6%
AG News, names87.0%88.5%89.0%89.0%
AG News, descriptions88.0%89.0%90.5%90.0%
CLINC150 (never seen by Kev)50.0%60.3%62.0%68.5%
MASSIVE (never seen)70.3%77.1%76.0%82.9%
Financial tweets (never seen)22.5%66.0%69.5%71.0%

The smallest Kev beat Jev on TREC: 94.6% against 89.0% with label names (39 to 11 on the questions only one got right, p < 0.001) and 96.4% against 93.6% with descriptions (p = 0.02). On the tasks it had never seen it fell apart. It answered "Financials" to 152 of 200 tweets and was far behind Jev on CLINC150 and MASSIVE as well (p < 0.001 on each). Kev-4B was close to the 9B everywhere except TREC with label names. Against Jev on the unseen tasks, the 4B was clearly behind only on CLINC150 (p < 0.001); on MASSIVE (p = 0.06) and the tweets (p = 0.15) the gaps were not significant.

Probabilities: Jev's ran high, Kev's ran low

Both models return a probability for every option, and the point of that is to decide which answers to trust without a person checking them. So I measured two things: how far the stated probability sat from actual accuracy, and what share of messages a threshold could accept while keeping accepted answers 95% correct.

Average top probability minus accuracy (Kev-9B / Jev)Share accepted at 95% accuracy (Kev-9B / Jev)
BANKING77−9.3 / +13.9 points61.0% / 0%
TREC, names−3.8 / −0.198.0% / 68.4%
TREC, descriptions−0.3 / +0.9100% / 96.6%
AG News, names+1.3 / +6.784.0% / 76.0%
AG News, descriptions+1.3 / +7.190.0% / 0%
CLINC150−10.3 / +15.242.7% / 57.3%
MASSIVE−16.7 / +6.457.7% / 74.9%
Financial tweets−16.6 / +15.823.0% / 0%

All of these use the highest option probability. Jev's response also has a separate confidence field, which TypeSafe describes as how concentrated the distribution is; it differed from the top probability on 67 of 154 BANKING77 answers and 190 of 400 CLINC150 answers. Because a reader is likely to set a threshold on the field named confidence, I recomputed the accepted shares with it: CLINC150 60.0%, MASSIVE 80.0%, TREC with names 67.2%, and the same 0% on the three tasks above, since it is also rounded to two decimals. The conclusions do not change.

A positive gap means the model said it was surer than it turned out to be. Jev was over-confident on six of the eight and close to exact on TREC. Kev-9B was close on the datasets it knew and too unsure on the new ones, as the last post found. The thresholds were searched on the same samples they are scored on, so the accepted shares describe these samples.

The zeros need explaining, because Jev's accuracy was not low on those tasks. Every probability Jev returned was rounded to two decimals, and a large share came back at exactly 1.00. On BANKING77, 64 answers were at 1.00 and 58 of them were right: 90.6%. Nothing ranks above 1.00, so that bucket cannot be split. Lowering the threshold added answers, but no threshold lifted the accepted set above 91.6% (reached at 0.95). So the share is 0%. AG News with descriptions was the same (154 at 1.00, 94.2% right, and no lower threshold did better), and so were the financial tweets (70 at 1.00, 90.0% right, best 90.0%). Kev returns unrounded probabilities, and its most confident answers were more often right.

On the two new tasks where Jev's confident answers held up, it accepted more than Kev: 57.3% of CLINC150 against 42.7%, and 74.9% of MASSIVE against 57.7%. Its probabilities also ranked right answers above wrong ones slightly better there (AUROC 0.92 and 0.90, against Kev-9B's 0.91 and 0.88). On every dataset Kev trained on, Kev's ranking was better.

So neither model's stated probability can be taken at face value on a new task. Kev's was too low by 10–17 points, so a rule like "accept at 95%" accepts too little. Jev's was too high by 6–16 points on most tasks, so the same rule accepts too much. Only a threshold set on your own labelled sample avoids both.

Questions no option answers

CLINC150's out-of-scope questions are the case where a router should hand the message to a person. Jev gave them an average top probability of 0.55, against 0.93 for in-scope questions, and the probability alone separated the two groups well (AUROC 0.94, against 0.88 for Kev-9B). At the threshold that kept accepted answers 95% correct, 3 of the 100 out-of-scope questions got through with Jev and 5 with Kev-9B. Jev did that while accepting 229 of the 400 messages; Kev accepted 171.

Speed and cost

From my server in Korea, Jev's median time per request was 219–227 ms. Most of that is distance. The response carries a server-side time header (x-envoy-upstream-service-time), and over 30 BANKING77 requests with 77 options its median was 53 ms (range 36–192). Kev-9B on one A100 in the same machine took 29–65 ms per request in the first post, with the slow end on BANKING77.

Jev charges $0.042 per million input tokens, and output is free. A BANKING77 request with 77 options was about 1,000 tokens, so 1,000 such messages cost about four cents. This whole comparison, 2,329 requests and 1.49 million tokens, cost $0.06. Kev's weights are free, but it needs a GPU; the 9B ran on an A100 here.

Which to use

  • Your task looks like something Kev trained on: Kev-9B was at least as accurate and its probabilities were more usable. Its README lists its training sources, so check them first.
  • Your task is new: Jev was more accurate on two of three unseen datasets and never clearly worse. Measure it on a few hundred of your own labelled messages before trusting either.
  • You plan to auto-accept confident answers: set the threshold on your own labels. If Jev's answers at 1.00 are not 95% correct on your data, no threshold will get you there. Check that bucket first.
  • Data cannot leave your machines: Kev runs locally; Jev is an API.

What this does not show

One Jev version (jev-1.13.0, pinned instead of jev-latest, which can move), calibration and paired tests mostly for Kev-9B, and eight conditions from six English datasets, all short texts with intent or topic labels. Jev's training data is not published, so the split into "trained" and "unseen" applies to Kev only. The latency figures include the network path from Korea. One TREC request failed with a 503 on the first pass and was retried once; nothing else failed.

Setup: Jev through POST https://api.typesafe.ai/v1/systemone with model: "jev-1.13.0" (the response reported the same ID), one request at a time from a server in Korea, 2026-09-29. Request body as for Kev: the message as state, one choice question with the dataset's instructions and one criterion per label (label description or null). Kev-9B, Kev-4B, Kev-0.8B, MoJev and the classifier rows are the results from the first and second posts, read from the same result files; rows are paired by position and the gold labels were checked to match. Samples: BANKING77 154 (seed 20260918), TREC test set 500, AG News 200, CLINC150 400 (300 in-scope plus 100 out-of-scope), MASSIVE English 175, financial-news tweets 200 (seed 20260926). Calibration and the accepted shares use the top-option probability, not Jev's confidence field; the shares recomputed on confidence are given in the text. The gap column is the mean top probability minus accuracy. The 95% share is the largest share of messages above any single threshold whose accepted answers are at least 95% correct. McNemar tests are exact and two-sided. Server-side time is the x-envoy-upstream-service-time response header over the first 30 BANKING77 messages. Cost from the usage.input_tokens in each response and the price on TypeSafe's models page. One TREC-with-descriptions request returned HTTP 503 and was retried once; the result file marks it.

Courses that go with this post

SOTAAZ course

Courses are sold one at a time. If you want all 13, there is a $199 lifetime bundle.

Stay Updated

Follow us for the latest posts and tutorials

Subscribe to Newsletter

Related Posts