Learn AI by Building

From your first dataset to production agents โ€” deep-dive series, hands-on notebooks, and experiments you can rerun yourself.

Paper of the Week

All Issues โ†’
Issue #6 ยท WeeklyOct 9, 2026

Paper of the Week #6 โ€” The Check Passed. What Did It Check?

This week: a Lean proof that does not match the paper it formalizes, a decision-model table whose test sets are in its training mix, labels that override definitions, memory that wins or loses depending on how much the model reads, NCCL symmetric memory gains that depend on payload size and GPU, a 40-49% token cut that needs lowercase and thinking off, and an AI prescribing pilot whose first phase has two physicians check every prescription.

Premium Series

Our Products

Courses and code built from what we measure here

Latest Posts

View All โ†’
How Far Can You Shrink a Decision Model? decider-4b from BF16 to Q2_K on llama.cpp

How Far Can You Shrink a Decision Model? decider-4b from BF16 to Q2_K on llama.cpp

We ran the open decision model decider-4b at five GGUF sizes, from BF16 (10.4 GiB of GPU memory) to Q2_K (4.4 GiB), on 500 TREC questions. Accuracy differences were too small to detect at every size. The probabilities moved: log-loss got slightly better at Q4_K_M and worse at Q8_0 (barely), Q3_K_M and Q2_K. Q2_K changed 13% of the answers and Q3_K_M cut the share of answers above 0.9 from 40% to 23%. Q4_K_M halved the memory; no accuracy difference was detected and its log-loss improved.

- Models & Algorithms
Read More
Can You Act on Microsoft-Decision-1's Probabilities? We Swapped the Labels on 500 Questions

Can You Act on Microsoft-Decision-1's Probabilities? We Swapped the Labels on 500 Questions

We sent 500 TREC questions to Microsoft-Decision-1 and Jev 1.13 through OpenRouter, and ran two local decision models, decider-4b and laya, with normal labels and with option names swapped against their definitions. Decision-1 followed the definitions (95.2% against 96.4%, not a difference by our rule) and at a 0.9 threshold handled 80% of the swapped cases on its own, 5 of those 400 wrong. laya followed the labels and was wrong on 110 of the 115 swapped cases it would have handled. The same request sent twice gave slightly different probabilities on Decision-1.

- Models & Algorithms
Read More
Calling Microsoft-Decision-1 on OpenRouter: Six Things We Hit on Day One

Calling Microsoft-Decision-1 on OpenRouter: Six Things We Hit on Day One

A short field guide from our first day with Microsoft-Decision-1 through OpenRouter: it uses the System One endpoint, not chat completions; choice options must be a map; confidence is not the top probability; the model ID carries a date; requests hit rate limits; and the same question cost us 3.5 to 4.8 times as much on Jev. Working Python included.

- Models & Algorithms
Read More
Does Codemode Work With Small Local Models? Qwen3.5-9B Got 28 of 30 Writing Code, 19 Calling Tools

Does Codemode Work With Small Local Models? Qwen3.5-9B Got 28 of 30 Writing Code, 19 Calling Tools

Armin Ronacher writes that codemode, where the model writes code that calls tools instead of calling them one at a time, does not yet work with smaller models. We tested it on 30 tasks against a mock issue tracker with Qwen3.5-4B, Qwen3.5-9B and Qwen3.8-27B in llama.cpp. Codemode got more tasks right at every size (22 vs 18, 28 vs 19, 30 vs 26); only the 9B difference passes our test (p = 0.012), and most of it comes from tool-calling replies cut off at our 2,048-token output limit. On tasks both modes got right, codemode's median tokens per task were 28-45% of tool calling's.

- Models & Algorithms
Read More
Speculative Decoding Speedups Depend on the Task: Prompt Lookup 6.2ร— on a File Edit, DFlash 0.7ร— on an Essay (Qwen3.5-9B, llama.cpp)

Speculative Decoding Speedups Depend on the Task: Prompt Lookup 6.2ร— on a File Edit, DFlash 0.7ร— on an Essay (Qwen3.5-9B, llama.cpp)

We measured three kinds of speculative decoding in llama.cpp on one Qwen3.5-9B file, alone on an A100: prompt n-gram lookup, the model's own MTP head, and a DFlash draft model. Rewriting a 130-line file ran 6.2ร— faster with n-gram lookup and 3.2ร— with DFlash. On a 500-word essay, MTP was 4% faster, n-gram lookup slightly slower and DFlash 0.69ร—. At temperature 0, n-gram and MTP output matched the plain run token for token; DFlash output diverged on three of four tasks.

- Models & Algorithms
Read More
How Much Context Fits on an 8 GB GPU? Qwen3.5-9B From 8K to 256K, Against Five Real Tasks

How Much Context Fits on an 8 GB GPU? Qwen3.5-9B From 8K to 256K, Against Five Real Tasks

We measured GPU memory for Qwen3.5-9B Q4_K_M at 8K to 256K context and counted the tokens of common tasks. On an 8 GB card one conversation gets 32K with the default KV cache and 64K with an 8-bit cache. A 20-turn coding chat reached 36,736 tokens, a 76-page paper 77,800, and coding agents send 648 to 15,970 tokens before you type a second message. Qwen3-8B, which keeps a cache in every layer, already needs 9.5 GiB at 32K.

- Models & Algorithms
Read More