Learn AI by Building
From your first dataset to production agents — deep-dive series, hands-on notebooks, and experiments you can rerun yourself.
Paper of the Week
All Issues →Paper of the Week #5 — Stale Caches, Burst FLOPS and Decision Models Going Local
This week: Context Language Models reuse a stale KV cache after editing their own context, two ways to run Jev-style decision models yourself, sustained GPU FLOPS 8 to 17 points below the burst number on four NVIDIA cards, and a map eval where always answering water scores 66%.
Premium Series
Our Products
Courses and code built from what we measure here
Video Courses
13 hands-on courses — quantization, diffusion, RAG, on-device AI, decision models. $199 lifetime bundle, first 3 lectures of every course free
LLM Quantization and Compression Hands-On
GPTQ, AWQ, GGUF, QLoRA — fit LLMs into the memory you have. The course behind our KV-cache measurements
Free Benchmark Code
The scripts and raw results behind the measurement posts, starting with the KV cache harness. Free with an account
Premium Series
140 deep-dive posts across 21 series, bilingual KO/EN, with production-ready code and notebooks
Latest Posts
View All →
Paper of the Week #5 — Stale Caches, Burst FLOPS and Decision Models Going Local
This week: Context Language Models reuse a stale KV cache after editing their own context, two ways to run Jev-style decision models yourself, sustained GPU FLOPS 8 to 17 points below the burst number on four NVIDIA cards, and a map eval where always answering water scores 66%.

GPT-2 with ReLU² Instead of GELU: 0.051 Lower Loss After 1B Tokens, From the Same Starting Weights
Three GPT-2 (124M) runs with ReLU² in the MLP ended at 3.623 mean validation loss after 1B FineWeb-Edu tokens, against 3.674 for GPT-2's GELU, at GPT-2's learning rate. Each pair started from identical weights, so the activation is the only difference. The gap clears the series' pre-registered bar (0.010) five times over and was below it at no checkpoint.

Pretraining Without Backprop, Rerun: What Happens to Dust When Backprop Gets Adam or More Epochs?
We reran Q Labs' Dust at 1M tokens with its own code. Under the paper's SGD setup our numbers matched the paper's, and with 20 backprop seeds the gap between Dust at 1,024 draws and backprop was under 0.02 (5.938 vs 5.941), smaller than the paper's 0.025. Backprop with the paper's Adam settings reached 5.384 in the same single epoch, and 16 epochs over the same 1M tokens reached 5.209 with SGD while using about 5% of Dust's compute. Our first Adam run came out worse than SGD; the paper's own settings fixed it, and we note what to check in your own runs.

Where Decision Models Go Wrong: Four Situations Behind the Sentences Most of Them Miss
We re-read the per-sentence results of up to 14 Jev-style decision models on six datasets and kept the 433 sentences at least half of them got wrong. Grouping those sentences and checking a few simple features showed four recurring situations: the question's shape and the dataset's rule disagree, categories overlap, the sentence contains another option's name, and no option fits. Examples, a full table of groups, and what to check in your own options.

Is EmbeddingGemma 2 Better Than EmbeddingGemma 1 for Korean Search? We Tested It on 177 of Our Own Posts
We searched our 177 published Korean/English post pairs with each post's description as the query. With Korean queries, EmbeddingGemma 1 ranked the right post first more often than EmbeddingGemma 2: 162 vs 153 posts in Korean (p = 0.022) and 164 vs 155 when the target was the English version (p = 0.035). English-to-English showed no difference. Pre-registered, text only.

Does laya Read Your Definitions or Your Labels? A Pre-Registered Rerun of the Label-Bias Test
We reran the 'labels override definitions' tests from arXiv 2610.02586 on two laya 0.3.7 checkpoints with our own TREC and AG News sets. Changing one line of prompt rendering made predictions identical whatever the labels were called. When labels contradicted the definitions, laya's TREC accuracy fell from 88.6% to 22.2%, and its top probability ranked right and wrong answers backwards (AUROC 0.36), the same direction the paper reports.