Learn AI by Building

From your first dataset to production agents — deep-dive series, hands-on notebooks, and experiments you can rerun yourself.

Paper of the Week

All Issues →
Issue #5 · WeeklyOct 8, 2026

Paper of the Week #5 — Stale Caches, Burst FLOPS and Decision Models Going Local

This week: Context Language Models reuse a stale KV cache after editing their own context, two ways to run Jev-style decision models yourself, sustained GPU FLOPS 8 to 17 points below the burst number on four NVIDIA cards, and a map eval where always answering water scores 66%.

Premium Series

Our Products

Courses and code built from what we measure here

Latest Posts

View All →
Paper of the Week #5 — Stale Caches, Burst FLOPS and Decision Models Going Local

Paper of the Week #5 — Stale Caches, Burst FLOPS and Decision Models Going Local

This week: Context Language Models reuse a stale KV cache after editing their own context, two ways to run Jev-style decision models yourself, sustained GPU FLOPS 8 to 17 points below the burst number on four NVIDIA cards, and a map eval where always answering water scores 66%.

- AI Research
Read More
GPT-2 with ReLU² Instead of GELU: 0.051 Lower Loss After 1B Tokens, From the Same Starting Weights

GPT-2 with ReLU² Instead of GELU: 0.051 Lower Loss After 1B Tokens, From the Same Starting Weights

Three GPT-2 (124M) runs with ReLU² in the MLP ended at 3.623 mean validation loss after 1B FineWeb-Edu tokens, against 3.674 for GPT-2's GELU, at GPT-2's learning rate. Each pair started from identical weights, so the activation is the only difference. The gap clears the series' pre-registered bar (0.010) five times over and was below it at no checkpoint.

- Models & Algorithms
Read More
Pretraining Without Backprop, Rerun: What Happens to Dust When Backprop Gets Adam or More Epochs?

Pretraining Without Backprop, Rerun: What Happens to Dust When Backprop Gets Adam or More Epochs?

We reran Q Labs' Dust at 1M tokens with its own code. Under the paper's SGD setup our numbers matched the paper's, and with 20 backprop seeds the gap between Dust at 1,024 draws and backprop was under 0.02 (5.938 vs 5.941), smaller than the paper's 0.025. Backprop with the paper's Adam settings reached 5.384 in the same single epoch, and 16 epochs over the same 1M tokens reached 5.209 with SGD while using about 5% of Dust's compute. Our first Adam run came out worse than SGD; the paper's own settings fixed it, and we note what to check in your own runs.

- Models & Algorithms
Read More
Where Decision Models Go Wrong: Four Situations Behind the Sentences Most of Them Miss

Where Decision Models Go Wrong: Four Situations Behind the Sentences Most of Them Miss

We re-read the per-sentence results of up to 14 Jev-style decision models on six datasets and kept the 433 sentences at least half of them got wrong. Grouping those sentences and checking a few simple features showed four recurring situations: the question's shape and the dataset's rule disagree, categories overlap, the sentence contains another option's name, and no option fits. Examples, a full table of groups, and what to check in your own options.

- Models & Algorithms
Read More
Is EmbeddingGemma 2 Better Than EmbeddingGemma 1 for Korean Search? We Tested It on 177 of Our Own Posts

Is EmbeddingGemma 2 Better Than EmbeddingGemma 1 for Korean Search? We Tested It on 177 of Our Own Posts

We searched our 177 published Korean/English post pairs with each post's description as the query. With Korean queries, EmbeddingGemma 1 ranked the right post first more often than EmbeddingGemma 2: 162 vs 153 posts in Korean (p = 0.022) and 164 vs 155 when the target was the English version (p = 0.035). English-to-English showed no difference. Pre-registered, text only.

- Models & Algorithms
Read More
Does laya Read Your Definitions or Your Labels? A Pre-Registered Rerun of the Label-Bias Test

Does laya Read Your Definitions or Your Labels? A Pre-Registered Rerun of the Label-Bias Test

We reran the 'labels override definitions' tests from arXiv 2610.02586 on two laya 0.3.7 checkpoints with our own TREC and AG News sets. Changing one line of prompt rendering made predictions identical whatever the labels were called. When labels contradicted the definitions, laya's TREC accuracy fell from 88.6% to 22.2%, and its top probability ranked right and wrong answers backwards (AUROC 0.36), the same direction the paper reports.

- Models & Algorithms
Read More