Learn AI by Building
From your first dataset to production agents โ deep-dive series, hands-on notebooks, and experiments you can rerun yourself.
Paper of the Week
All Issues โPaper of the Week #4 โ A Memory of Procedures, or a Memory of Examples?
Designer-RSI grows a natural-language skill bank from user traffic and lifts execution success from 72.7% to 99.3% with no weight updates. I built the narrow version on a task with human labels: 40 rules distilled from the model's own mistakes fixed 4 items and broke 5. Retrieving five raw examples fixed 19 and broke none.
Premium Series
Our Products
Courses and code built from what we measure here
Video Courses
13 hands-on courses โ quantization, diffusion, RAG, on-device AI, decision models. $199 lifetime bundle, first 3 lectures of every course free
LLM Quantization and Compression Hands-On
GPTQ, AWQ, GGUF, QLoRA โ fit LLMs into the memory you have. The course behind our KV-cache measurements
Free Benchmark Code
The scripts and raw results behind the measurement posts, starting with the KV cache harness. Free with an account
Premium Series
140 deep-dive posts across 21 series, bilingual KO/EN, with production-ready code and notebooks
Latest Posts
View All โ
Strata vs llama.cpp on the Same Qwen3.8-Flash-Next File: No Score Difference Found, Faster Within 12 GiB
Same IQ2_XS file in both engines: GSM8K 93.3% vs 93.0%, HumanEval 157/164 each, no vision gap showed up on 100 COCO images. Held to 12 GiB, Strata (draft decoding on) decoded 62.9 tok/s to llama.cpp's 23.4.

llama.cpp KV Cache Quantization: Which -ctk and -ctv to Use, Measured on the Current Build
Keep f16 if the KV cache fits. If not, set -ctk q8_0 -ctv q8_0: it saved 4.2 GiB at 64K but decoded at 56% of f16's speed there.

What Are You Deciding, and Which Mistakes Are Expensive? Chapter 1 of a Book on Building Decision Systems
Free chapter: before training anything, count what each intent can tell you, find the intents likely to be confused, and write down what each mistake costs. The code runs in under a second.

Which Mistakes Are Expensive? One Routing System Re-scored at Seven Prices for a Wrong Answer
Re-scoring 3,080 BANKING77 messages: the price of a wrong answer moved the threshold far more than the choice of model, and decision models paid only when a wrong answer cost two hand-offs or less.

How Far Does a Simple Classifier Get? Chapter 2 of a Book on Building Decision Systems
Free sample chapter: split BANKING77 before training, then two CPU classifiers reach 91.2% and 92.9% on the test set, with code that runs in about a minute.

Which Messages Should Go to a Person? A Classifier, a Decision Model and a Hand-Off, Measured
On BANKING77 a decision model after a classifier saved no hand-offs. On CLINC150 unknown questions broke the thresholds; adding 250 to validation halved the leaks.