Learn AI by Building
From your first dataset to production agents โ deep-dive series, hands-on notebooks, and experiments you can rerun yourself.
Paper of the Week
All Issues โPaper of the Week #4 โ A Memory of Procedures, or a Memory of Examples?
Designer-RSI grows a natural-language skill bank from user traffic and lifts execution success from 72.7% to 99.3% with no weight updates. I built the narrow version on a task with human labels: 40 rules distilled from the model's own mistakes fixed 4 items and broke 5. Retrieving five raw examples fixed 19 and broke none.
Premium Series
Our Products
Courses and code built from what we measure here
Video Courses
13 hands-on courses โ quantization, diffusion, RAG, on-device AI, decision models. $199 lifetime bundle, first 3 lectures of every course free
LLM Quantization and Compression Hands-On
GPTQ, AWQ, GGUF, QLoRA โ fit LLMs into the memory you have. The course behind our KV-cache measurements
Free Benchmark Code
The scripts and raw results behind the measurement posts, starting with the KV cache harness. Free with an account
Premium Series
140 deep-dive posts across 21 series, bilingual KO/EN, with production-ready code and notebooks
Latest Posts
View All โ
llama.cpp KV Cache Quantization: Which -ctk and -ctv to Use, Measured on the Current Build
Keep f16 if the KV cache fits. If not, set -ctk q8_0 -ctv q8_0: it saved 4.2 GiB at 64K but decoded at 56% of f16's speed there.

What Are You Deciding, and Which Mistakes Are Expensive? Chapter 1 of a Book on Building Decision Systems
Free chapter: before training anything, count what each intent can tell you, find the intents likely to be confused, and write down what each mistake costs. The code runs in under a second.

Which Mistakes Are Expensive? One Routing System Re-scored at Seven Prices for a Wrong Answer
Re-scoring 3,080 BANKING77 messages: the price of a wrong answer moved the threshold far more than the choice of model, and decision models paid only when a wrong answer cost two hand-offs or less.

How Far Does a Simple Classifier Get? Chapter 2 of a Book on Building Decision Systems
Free sample chapter: split BANKING77 before training, then two CPU classifiers reach 91.2% and 92.9% on the test set, with code that runs in about a minute.

Which Messages Should Go to a Person? A Classifier, a Decision Model and a Hand-Off, Measured
On BANKING77 a decision model after a classifier saved no hand-offs. On CLINC150 unknown questions broke the thresholds; adding 250 to validation halved the leaks.

Classifier, LLM or Decision Model? A Measured Guide to Text Classification
One path through every text-classification measurement on this blog: on the same 154 banking messages, a CPU classifier scored 90.3%, an LLM with five retrieved examples 94.8%, and Jev 76.0%. Which to use, and when.