Learn AI by Building

From your first dataset to production agents โ€” deep-dive series, hands-on notebooks, and experiments you can rerun yourself.

Paper of the Week

All Issues โ†’
Issue #4 ยท WeeklySep 30, 2026

Paper of the Week #4 โ€” A Memory of Procedures, or a Memory of Examples?

Designer-RSI grows a natural-language skill bank from user traffic and lifts execution success from 72.7% to 99.3% with no weight updates. I built the narrow version on a task with human labels: 40 rules distilled from the model's own mistakes fixed 4 items and broke 5. Retrieving five raw examples fixed 19 and broke none.

Premium Series

Our Products

Courses and code built from what we measure here

Latest Posts

View All โ†’
llama.cpp KV Cache Quantization: Which -ctk and -ctv to Use, Measured on the Current Build

llama.cpp KV Cache Quantization: Which -ctk and -ctv to Use, Measured on the Current Build

Keep f16 if the KV cache fits. If not, set -ctk q8_0 -ctv q8_0: it saved 4.2 GiB at 64K but decoded at 56% of f16's speed there.

- Models & Algorithms
Read More
What Are You Deciding, and Which Mistakes Are Expensive? Chapter 1 of a Book on Building Decision Systems

What Are You Deciding, and Which Mistakes Are Expensive? Chapter 1 of a Book on Building Decision Systems

Free chapter: before training anything, count what each intent can tell you, find the intents likely to be confused, and write down what each mistake costs. The code runs in under a second.

- Models & Algorithms
Read More
Which Mistakes Are Expensive? One Routing System Re-scored at Seven Prices for a Wrong Answer

Which Mistakes Are Expensive? One Routing System Re-scored at Seven Prices for a Wrong Answer

Re-scoring 3,080 BANKING77 messages: the price of a wrong answer moved the threshold far more than the choice of model, and decision models paid only when a wrong answer cost two hand-offs or less.

- Models & Algorithms
Read More
How Far Does a Simple Classifier Get? Chapter 2 of a Book on Building Decision Systems

How Far Does a Simple Classifier Get? Chapter 2 of a Book on Building Decision Systems

Free sample chapter: split BANKING77 before training, then two CPU classifiers reach 91.2% and 92.9% on the test set, with code that runs in about a minute.

- Models & Algorithms
Read More
Which Messages Should Go to a Person? A Classifier, a Decision Model and a Hand-Off, Measured

Which Messages Should Go to a Person? A Classifier, a Decision Model and a Hand-Off, Measured

On BANKING77 a decision model after a classifier saved no hand-offs. On CLINC150 unknown questions broke the thresholds; adding 250 to validation halved the leaks.

- Models & Algorithms
Read More
Classifier, LLM or Decision Model? A Measured Guide to Text Classification

Classifier, LLM or Decision Model? A Measured Guide to Text Classification

One path through every text-classification measurement on this blog: on the same 154 banking messages, a CPU classifier scored 90.3%, an LLM with five retrieved examples 94.8%, and Jev 76.0%. Which to use, and when.

- Models & Algorithms
Read More