Learn AI by Building

From your first dataset to production agents β€” deep-dive series, hands-on notebooks, and experiments you can rerun yourself.

Paper of the Week

All Issues β†’
Issue #6 Β· WeeklyOct 9, 2026

Paper of the Week #6 β€” The Check Passed. What Did It Check?

This week: a Lean proof that does not match the paper it formalizes, a decision-model table whose test sets are in its training mix, labels that override definitions, memory that wins or loses depending on how much the model reads, NCCL symmetric memory gains that depend on payload size and GPU, a 40-49% token cut that needs lowercase and thinking off, and an AI prescribing pilot whose first phase has two physicians check every prescription.

Premium Series

Our Products

Courses and code built from what we measure here

Latest Posts

View All β†’
How Much Context Fits on an 8 GB GPU? Qwen3.5-9B From 8K to 256K, Against Five Real Tasks

How Much Context Fits on an 8 GB GPU? Qwen3.5-9B From 8K to 256K, Against Five Real Tasks

We measured GPU memory for Qwen3.5-9B Q4_K_M at 8K to 256K context and counted the tokens of common tasks. On an 8 GB card one conversation gets 32K with the default KV cache and 64K with an 8-bit cache. A 20-turn coding chat reached 36,736 tokens, a 76-page paper 77,800, and coding agents send 648 to 15,970 tokens before you type a second message. Qwen3-8B, which keeps a cache in every layer, already needs 9.5 GiB at 32K.

- Models & Algorithms
Read More
Paper of the Week #6 β€” The Check Passed. What Did It Check?

Paper of the Week #6 β€” The Check Passed. What Did It Check?

This week: a Lean proof that does not match the paper it formalizes, a decision-model table whose test sets are in its training mix, labels that override definitions, memory that wins or loses depending on how much the model reads, NCCL symmetric memory gains that depend on payload size and GPU, a 40-49% token cut that needs lowercase and thinking off, and an AI prescribing pilot whose first phase has two physicians check every prescription.

- AI Research
Read More
We Trained Our Own Decision Model with Unsloth: 75% on BANKING77, 56% When BANKING77 Is Left Out of Training

We Trained Our Own Decision Model with Unsloth: 75% on BANKING77, 56% When BANKING77 Is Left Out of Training

We followed Unsloth's recipe for turning Qwen3.5-0.8B into a decision model (4.15 GB peak memory, about 20 minutes per run on a shared GPU). BANKING77 and CLINC150 matched Unsloth's table (75.4% and 75.0%), but those test rows come from datasets in the training mix. With the two datasets removed from training, the same 500 rows scored 56.5% and 55.9%. typed-decisions came out 11.8 points below the table; the recipe's decontamination step drops 850 of its 1,200 training cases, and in a post hoc run that kept them it reached 74.3%.

- Models & Algorithms
Read More
Paper of the Week #5 β€” Stale Caches, Burst FLOPS and Decision Models Going Local

Paper of the Week #5 β€” Stale Caches, Burst FLOPS and Decision Models Going Local

This week: Context Language Models reuse a stale KV cache after editing their own context, two ways to run Jev-style decision models yourself, sustained GPU FLOPS 8 to 17 points below the burst number on four NVIDIA cards, and a map eval where always answering water scores 66%.

- AI Research
Read More
GPT-2 with ReLUΒ² Instead of GELU: 0.051 Lower Loss After 1B Tokens, From the Same Starting Weights

GPT-2 with ReLUΒ² Instead of GELU: 0.051 Lower Loss After 1B Tokens, From the Same Starting Weights

Three GPT-2 (124M) runs with ReLUΒ² in the MLP ended at 3.623 mean validation loss after 1B FineWeb-Edu tokens, against 3.674 for GPT-2's GELU, at GPT-2's learning rate. Each pair started from identical weights, so the activation is the only difference. The gap clears the series' pre-registered bar (0.010) five times over and was below it at no checkpoint.

- Models & Algorithms
Read More
Pretraining Without Backprop, Rerun: What Happens to Dust When Backprop Gets Adam or More Epochs?

Pretraining Without Backprop, Rerun: What Happens to Dust When Backprop Gets Adam or More Epochs?

We reran Q Labs' Dust at 1M tokens with its own code. Under the paper's SGD setup our numbers matched the paper's, and with 20 backprop seeds the gap between Dust at 1,024 draws and backprop was under 0.02 (5.938 vs 5.941), smaller than the paper's 0.025. Backprop with the paper's Adam settings reached 5.384 in the same single epoch, and 16 epochs over the same 1M tokens reached 5.209 with SGD while using about 5% of Dust's compute. Our first Adam run came out worse than SGD; the paper's own settings fixed it, and we note what to check in your own runs.

- Models & Algorithms
Read More