Models & Algorithms••KR

Jeff vs Jev on the Same Questions: Overall Scores and a 26-Option Limit

On Jeff's own 4,599 questions, Jev scored 85.7 and Jeff-2B 83.0. In my tests neither Jeff model picked an option past the 26th.

Jeff vs Jev on the Same Questions: Overall Scores and a 26-Option Limit

Jeff vs Jev on the Same Questions: Overall Scores and a 26-Option Limit

Jeff is a new pair of open decision models, fine-tuned from Qwen3.5-0.8B and Qwen3.5-2B, that accept the same request format as Jev. Its README puts Jeff-2B just ahead of Jev, 83.1 to 83.0 over five public benchmarks, and marks Jeff as the winner of that row. It is also careful to say that Jev's figure "was measured on a different sample of the same benchmarks".

I have Jev API access and the same harness I used for Kev vs Jev, so I rebuilt Jeff's benchmark with Jeff's own script and ran Jev, Jeff-0.8B, Jeff-2B and Kev-9B on the same 4,599 questions. Then I ran Jeff on the six datasets from my earlier posts.

The short answer

  • On the same questions, Jev was ahead: 85.7 against Jeff-2B's 83.0 (p < 0.001, paired). Jeff-2B was clearly better on Financial PhraseBank and RAGTruth, and clearly worse on BBH, JudgeBench and WinoGrande.
  • The overall score depends on how the five benchmarks are combined. Jeff's "Overall" weights them by question count, a common choice, and two of them make up 54% of the questions. A plain mean of the five gives 79.2 against 85.5 on my run.
  • In my tests Jeff never chose an option past the 26th. On intent tasks with 60 to 150 options it scored 9–36%, while Kev-9B and Jev scored 62–83%. Reversing the option list moved the failures with it.

Reproducing Jeff's benchmark

Jeff's repository builds its benchmark with a script, jeff-panel, that downloads pinned revisions of five public datasets and draws a fixed-seed sample: 750 BBH questions (15 tasks, 50 each), 999 Financial PhraseBank sentences, 350 JudgeBench pairs, 1,500 RAGTruth responses and 1,000 WinoGrande sentences. I ran it unchanged. Before comparing anything, I checked that Jeff scored on my rebuild what its README reports:

Jeff-0.8B, READMEJeff-0.8B, my runJeff-2B, READMEJeff-2B, my run
BBH64.064.068.067.6
Financial PhraseBank96.496.696.396.3
JudgeBench62.662.364.664.3
RAGTruth86.186.288.988.8
WinoGrande68.669.279.079.2
Overall79.179.383.183.0

Every figure is within 0.6 points, which is consistent with the questions and the serving setup being the same as Jeff's.

Jev on the same 4,599 questions

Dot chart, same 4,599 questions for all four models. BBH: Jeff-2B 67.6, Jev 90.5. Financial PhraseBank: Jeff-2B 96.3, Jev 85.3. JudgeBench: Jeff-2B 64.3, Jev 79.4. RAGTruth: Jeff-2B 88.8, Jev 81.8. WinoGrande: Jeff-2B 79.2, Jev 90.6. Jeff's weighted overall: Jeff-2B 83.0, Jev 85.7. Plain mean of the five: Jeff-2B 79.2, Jev 85.5. Jeff-0.8B is a hollow marker and Kev-9B a small dot.
Jeff-0.8BJeff-2BKev-9BJev 1.13Only one got it right (Jeff-2B / Jev)
BBH (750)64.067.665.790.526 / 198, p < 0.001
Financial PhraseBank (999)96.696.393.785.3132 / 22, p < 0.001
JudgeBench (350)62.364.362.379.431 / 84, p < 0.001
RAGTruth (1,500)86.288.868.181.8192 / 87, p < 0.001
WinoGrande (1,000)69.279.272.390.650 / 164, p < 0.001
Jeff's overall (weighted by count)79.383.073.885.7431 / 555, p < 0.001
Plain mean of the five75.779.272.485.5

On these questions Jev scored 85.7, not the 83.0 in Jeff's README. The biggest change was Financial PhraseBank: 85.3 here against the 77.0 Jeff quotes. Jeff's panel uses the "all annotators agree" subset of Financial PhraseBank, the sentences on which every annotator gave the same sentiment, which is the easiest version of that dataset. That is a fair choice, but it means the 77.0 and the 85.3 are not the same test.

On the same questions Jeff-2B beat Jev on Financial PhraseBank by 11 points and on RAGTruth by 7. Jev was ahead on the other three: by 11 points on WinoGrande, 15 on JudgeBench and 23 on BBH. On BBH the gap was 23 points, and it was widest on the multi-step tasks: on reading a table of penguins Jev got 49 of 50 and Jeff-2B 20; on tracking shuffled objects, 47 against 24.

Training data matters for reading these numbers. Jeff's data sources list the training splits of RAGTruth and WinoGrande and a financial-news sentiment dataset of tweets, so three of the five panel benchmarks are close to tasks Jeff trained on (the panel items themselves were filtered out of training). TypeSafe does not publish what Jev was trained on, so I can't say the same for Jev either way.

Kev-9B, trained on none of these five, came in below both Jeff models overall. That fits what the Kev posts found: it does well on tasks shaped like its training data, and these reasoning and grounding benchmarks are not.

How benchmark weighting affects the overall score

Jeff's "Overall" is the share of all 4,599 questions answered correctly, so each benchmark counts in proportion to its size. That is a standard way to aggregate. The alternative, a plain mean of the five benchmark scores, gives each benchmark equal weight. The two can disagree here because the benchmarks differ in size: RAGTruth has 1,500 questions and Financial PhraseBank 999, 54% of the total, and those are Jeff's two strongest benchmarks.

Jev's published per-benchmark figures, weighted the same way, give exactly 83.0; averaged plainly, they give 83.6. Jeff-2B's published figures give 83.1 weighted and 79.4 averaged. So on the published figures the order depends on the aggregation. On the same questions, Jev is ahead under both.

Calibration: good overall, uneven by benchmark

Jeff's model card reports a calibration error of 0.028 for the 2B and 0.049 for the 0.8B. Pooling all 4,599 questions I got 0.028 and 0.046, so the card's figures appear to be pooled ones. By benchmark, Jeff-2B's error ranged from 1.2 points on RAGTruth to 14.6 on JudgeBench, where its average top probability was 13.3 points above its accuracy. Jev's ranged from 2.5 to 4.2 points on every benchmark, and 1.5 pooled.

A pooled number can look good while one benchmark is badly off, because errors in opposite directions cancel: Jeff-2B was over-confident on JudgeBench (+13.3 points) and WinoGrande (+6.9) and under-confident on Financial PhraseBank (−5.1). Using the same rule as the earlier posts, a threshold set so that accepted answers stay 95% correct would have accepted 54.4% of all questions for Jeff-2B and 63.0% for Jev. On JudgeBench it would have accepted 0.6% for Jeff-2B and 29.1% for Jev.

Off the benchmark: options past the 26th

Then I ran both Jeff models on the eight conditions from Kev vs Jev, same rows, same label names:

Jeff-0.8BJeff-2BKev-9BJev 1.13
BANKING77 (77 options)23.4%16.2%82.5%76.0%
TREC, names (6)72.0%63.0%93.8%89.0%
TREC, descriptions (6)76.6%80.6%96.6%93.6%
AG News, names (4)88.0%88.5%89.0%89.0%
AG News, descriptions (4)88.0%90.0%90.5%90.0%
CLINC150 (150 + out-of-scope)10.8%9.2%62.0%68.5%
MASSIVE (60)36.0%34.9%76.0%82.9%
Financial tweets, topics (20)53.0%63.5%69.5%71.0%

CLINC150's 100 out-of-scope questions have no correct option, so every model gets them wrong and 75% is the ceiling there.

The drop was on the three tasks with more than 26 options, and the failures showed a strong option-position pattern. Jeff labels options with letters, A to Z, then AA, AB and so on, and reads its answer as the probability of each label. In all 1,458 answers on those three tasks, neither Jeff model picked an option labelled past Z. When the right option was among the first 26, Jeff-0.8B got 80% right, about the same as Kev-9B (79%). Jeff-2B got 69%; the smaller model was also ahead on BANKING77 (23.4% against 16.2%), and I did not find out why. When it was 27th or later, each Jeff model got 0 of 452.

Grouped bar chart, accuracy by the position of the right option, pooled over BANKING77, CLINC150 and MASSIVE. Options 1–26: Jeff-0.8B 80%, Jeff-2B 69%, Kev-9B 79%, Jev 84%. Options 27–52, 53–78, 79–104 and 105–150: Jeff-0.8B and Jeff-2B 0% in every group; Kev-9B 73–88%; Jev 80–95%.

To check that position, not the questions, was the cause, I reversed the order of BANKING77's 77 options and ran Jeff-2B again. The 102 messages it had got 0 of now scored 25; the ones whose answers had moved past Z now scored 0. The right option past Z got a median probability of 0.0004. Jeff's README says a choice can have up to 255 options and that "the options can be anything: support queues, user intents, moderation labels, voice commands, game moves." The server accepts 255, but the models did not use any beyond the 26th here. MASSIVE is in Jeff's training data, and it still scored 35–36% there, because 102 of the 175 right answers sat past Z.

If you use Jeff, keep each question to 26 options or fewer, for example by shortlisting candidates in code first. On AG News both Jeff models were within four articles of Jev (p ≥ 0.29). On the six-option TREC task they trailed Jev by 13 to 26 points, so option position does not explain every gap.

Which to use

  • Financial sentiment or checking a response against its source (RAGTruth-style): Jeff-2B was better than Jev on the same questions and runs locally.
  • Reasoning, judging between two answers, pronoun resolution: Jev was 11 to 23 points ahead.
  • More than 26 options: not Jeff, or shortlist to 26 first. Kev and Jev handled 150.

What this does not show

Jev is one version (jev-1.13.0). Jeff's panel is Jeff's own choice of benchmarks and samples; I kept it as is so the numbers compare with its README, and I did not run JevBench's hard tier or Jeff's game tests. Jeff was trained on the training splits of RAGTruth and WinoGrande, Kev on none of the five, and TypeSafe does not publish what Jev was trained on. The 26-option finding is from three intent datasets and one reversal test on one of them. I did not test Jeff-Gemma4-E2B.

Setup: Jeff repository firelex/jeff at 985797f; weights mstrasser/Jeff-Qwen3.5-0.8B 9878256 and mstrasser/Jeff-Qwen3.5-2B f2d993c, served with the repository's jeff-serve (PyTorch, bf16, torch 2.14.0) on one A100 80GB PCIe, 127.0.0.1 only, model name jeff-latest. Panel: jeff-panel unchanged, seed 20260926, 4,599 rows, none dropped for length, SHA-256 46b6fb82…. Kev-9B as in the earlier posts, served with kev.serve. Jev: POST https://api.typesafe.ai/v1/systemone with model: "jev-1.13.0", four requests at a time, no errors or retries. Every model got the panel's own state, instructions and criteria; yes/no questions (RAGTruth) are scored as "yes" at a probability of 0.5 or more. Calibration error is ECE over 10 bins of the top-option probability (for yes/no, max(p, 1 − p)); McNemar tests are exact and two-sided. The eight-condition runs use the same files, label names and instructions as the Kev posts. The reversal test used the same 154 BANKING77 messages with the option list reversed, and every response's model field was jeff-qwen3.5-2b. Speed: Jeff answered the panel in a median of 45 ms per question on the A100 over local HTTP, for both sizes (Jeff's README reports 22–24 ms on an RTX PRO 6000, a setup I did not reproduce). The 4,599 Jev requests used 3.39 million input tokens, $0.14 at list price. Measured 2026-09-29.

Courses that go with this post

SOTAAZ course

Courses are sold one at a time. If you want all 13, there is a $199 lifetime bundle.

Stay Updated

Follow us for the latest posts and tutorials

Subscribe to Newsletter

Related Posts

Kev vs MoJev: Two More Open Jev Alternatives, Accuracy and Confidence Measured
Models & Algorithms

Kev vs MoJev: Two More Open Jev Alternatives, Accuracy and Confidence Measured

Kev (0.8B to 9B) and MoJev (0.85B) both speak Jev's API. On BANKING77, TREC and AG News with human labels, Kev's probabilities were usable: at 95% accuracy it could auto-accept 53–61% of BANKING77, where laya managed 0%. But Kev was trained on these three datasets, and a small classifier trained on the same data still beat it on BANKING77. MoJev reports 0.79% calibration error; here it was 6–22%, with probabilities too low.

Can You Trust a Decision Model's Confidence? Repeated Answers and Probabilities, Measured
Models & Algorithms

Can You Trust a Decision Model's Confidence? Repeated Answers and Probabilities, Measured

Decision models sell a probability with every answer, so you can act on the sure ones and escalate the rest. I checked both ways of judging 'sure' on human-labelled data. Across 8 noise draws, 77–80% of openjev's wrong answers were unanimous. Probabilities did better, but how useful they were varied sharply by task: laya could auto-accept 88% of TREC at 95% accuracy and 0% of BANKING77. The new CLM-8B stayed near chance on all three datasets when given label names.

Jev in 10 Minutes: What a Fast Decision Model Is, and When It Is Worth Using
Models & Algorithms

Jev in 10 Minutes: What a Fast Decision Model Is, and When It Is Worth Using

Jev answers a question with a probability for each option instead of writing text. What that is for, how open projects do the same thing, and what three rounds of measurement showed: with a handful of options a 421M model reached 86.6% in 23 ms, with 77 a small trained classifier on a CPU still led at 90%.