Learn AI by Building
From your first dataset to production agents β deep-dive series, hands-on notebooks, and experiments you can rerun yourself.
Paper of the Week
All Issues βPaper of the Week #4 β A Memory of Procedures, or a Memory of Examples?
Designer-RSI grows a natural-language skill bank from user traffic and lifts execution success from 72.7% to 99.3% with no weight updates. I built the narrow version on a task with human labels: 40 rules distilled from the model's own mistakes fixed 4 items and broke 5. Retrieving five raw examples fixed 19 and broke none.
Premium Series
Our Products
Courses and code built from what we measure here
Video Courses
13 hands-on courses β quantization, diffusion, RAG, on-device AI, decision models. $199 lifetime bundle, first 3 lectures of every course free
LLM Quantization and Compression Hands-On
GPTQ, AWQ, GGUF, QLoRA β fit LLMs into the memory you have. The course behind our KV-cache measurements
Free Benchmark Code
The scripts and raw results behind the measurement posts, starting with the KV cache harness. Free with an account
Premium Series
140 deep-dive posts across 21 series, bilingual KO/EN, with production-ready code and notebooks
Latest Posts
View All β
Is EmbeddingGemma 2 Better Than EmbeddingGemma 1 for Korean Search? We Tested It on 177 of Our Own Posts
We searched our 177 published Korean/English post pairs with each post's description as the query. With Korean queries, EmbeddingGemma 1 ranked the right post first more often than EmbeddingGemma 2: 162 vs 153 posts in Korean (p = 0.022) and 164 vs 155 when the target was the English version (p = 0.035). English-to-English showed no difference. Pre-registered, text only.

Does laya Read Your Definitions or Your Labels? A Pre-Registered Rerun of the Label-Bias Test
We reran the 'labels override definitions' tests from arXiv 2610.02586 on two laya 0.3.7 checkpoints with our own TREC and AG News sets. Changing one line of prompt rendering made predictions identical whatever the labels were called. When labels contradicted the definitions, laya's TREC accuracy fell from 88.6% to 22.2%, and its top probability ranked right and wrong answers backwards (AUROC 0.36), the same direction the paper reports.

GPT-2 with Parameter-Free RMSNorm: 0.038 Higher Loss After 1B Tokens; QK Norm Won Back 0.012, Barely Past the Bar
Swapping all of GPT-2's LayerNorms for parameter-free RMSNorm raised the mean validation loss after 1B FineWeb-Edu tokens from 3.674 to 3.712, three runs each at GPT-2's learning rate. Adding QK norm on top brought it to 3.701, a 0.012 gain that clears this series' pre-registered bar (0.011) by 0.001. Throughput differences stayed within run-to-run variation, and the logged gradient norms showed no instability to fix.

GPT-2 with RoPE Instead of Learned Position Embeddings: 0.122 Lower Loss After 1B Tokens, 5% Slower
Three GPT-2 (124M) runs with rotary position embeddings ended at 3.552 mean validation loss after 1B FineWeb-Edu tokens, against 3.674 for three runs with GPT-2's learned position embeddings. The gap clears this series' pre-registered bar (0.012) by ten times, at GPT-2's learning rate. RoPE ran 5.3% slower on this code.

Same GPT-2, Same Data, Three Seeds: The Final Loss Spread 0.018 Nats
Three GPT-2 (124M) runs on 1B FineWeb-Edu tokens that differed only in the random seed ended at 3.668, 3.669 and 3.686 validation loss (sd 0.010). Seed 0 run again with the same command landed 0.0035 away. With three seeds per side, a change has to beat 0.016 nats before this series calls it real.

Karpathy's βLet's Reproduce GPT-2β in One Read: The Four Parts, Their Numbers, and What Changed Since
A condensed walk through the 4-hour video: building GPT-2 (124M), taking a step from 1,000 ms to 90 ms on one A100, the GPT-3 training settings, and a 10B-token run that beat OpenAI's checkpoint on HellaSwag. Plus one check we ran: nanoGPT's exact GELU moves logits by up to 5 against GPT-2's.