Learn AI by Building

From your first dataset to production agents โ€” deep-dive series, hands-on notebooks, and experiments you can rerun yourself.

Paper of the Week

All Issues โ†’
Issue #4 ยท WeeklySep 30, 2026

Paper of the Week #4 โ€” A Memory of Procedures, or a Memory of Examples?

Designer-RSI grows a natural-language skill bank from user traffic and lifts execution success from 72.7% to 99.3% with no weight updates. I built the narrow version on a task with human labels: 40 rules distilled from the model's own mistakes fixed 4 items and broke 5. Retrieving five raw examples fixed 19 and broke none.

Premium Series

Our Products

Courses and starter kits built from what we measure here

Starter Kits

View All โ†’

Practice notebooks, interview questions, and project solutions โ€” ready to download.

Browse Starter Kits

Latest Posts

View All โ†’
Classifier, LLM or Decision Model? A Measured Guide to Text Classification

Classifier, LLM or Decision Model? A Measured Guide to Text Classification

One path through every text-classification measurement on this blog: on the same 154 banking messages, a CPU classifier scored 90.3%, an LLM with five retrieved examples 94.8%, and Jev 76.0%. Which to use, and when.

- Models & Algorithms
Read More
Ollama's Decision Models on the Same Questions as Jev: Nimble and Tev1, Measured

Ollama's Decision Models on the Same Questions as Jev: Nimble and Tev1, Measured

On the same questions as Jev, Ollama's Nimble 9B scored 95.6% on TREC (Jev 89.0%) but 76.0 against 85.7 on a 4,599-question reasoning-heavy panel.

- Models & Algorithms
Read More
Paper of the Week #4 โ€” A Memory of Procedures, or a Memory of Examples?

Paper of the Week #4 โ€” A Memory of Procedures, or a Memory of Examples?

Designer-RSI grows a natural-language skill bank from user traffic and lifts execution success from 72.7% to 99.3% with no weight updates. I built the narrow version on a task with human labels: 40 rules distilled from the model's own mistakes fixed 4 items and broke 5. Retrieving five raw examples fixed 19 and broke none.

- AI Research
Read More
Jeff vs Jev on the Same Questions: Overall Scores and a 26-Option Limit (Fixed in v1.1)

Jeff vs Jev on the Same Questions: Overall Scores and a 26-Option Limit (Fixed in v1.1)

On Jeff's own 4,599 questions, Jev scored 85.7 and Jeff-2B 83.0. Jeff v1.0 never picked an option past the 26th in my tests; v1.1 fixes that, remeasured.

- Models & Algorithms
Read More
Kev vs Jev: Kev 0.8B to 9B Benchmarked for Accuracy and Calibration

Kev vs Jev: Kev 0.8B to 9B Benchmarked for Accuracy and Calibration

Kev (0.8B, 4B, 9B) and Jev 1.13 on the same 1,629 labelled messages (2,329 requests). On the three datasets Kev trained on, Kev-9B was as accurate or more (TREC 93.8% vs 89.0%). On three it never saw, Jev was ahead on two (CLINC150 68.5% vs 62.0%, MASSIVE 82.9% vs 76.0%). Jev's probabilities ran high and come rounded to two decimals, so on BANKING77 no threshold reached 95% accuracy.

- Models & Algorithms
Read More
How to Run Qwen Locally: Which File Fits Your GPU (8, 12, 16 or 24 GB), Measured

How to Run Qwen Locally: Which File Fits Your GPU (8, 12, 16 or 24 GB), Measured

Which Qwen GGUF fits an 8, 12, 16 or 24 GB GPU, with measured memory: 9B Q4_K_M 5.8 GiB, 27B Q3_K_XL 13.1 GiB, 35B-A3B 7.3 GiB with 30 layers' experts on the CPU.

- Models & Algorithms
Read More