Learn AI by Building
From your first dataset to production agents β deep-dive series, hands-on notebooks, and experiments you can rerun yourself.
Paper of the Week
All Issues βPaper of the Week #4 β A Memory of Procedures, or a Memory of Examples?
Designer-RSI grows a natural-language skill bank from user traffic and lifts execution success from 72.7% to 99.3% with no weight updates. I built the narrow version on a task with human labels: 40 rules distilled from the model's own mistakes fixed 4 items and broke 5. Retrieving five raw examples fixed 19 and broke none.
Premium Series
Our Products
Courses and code built from what we measure here
Video Courses
13 hands-on courses β quantization, diffusion, RAG, on-device AI, decision models. $199 lifetime bundle, first 3 lectures of every course free
LLM Quantization and Compression Hands-On
GPTQ, AWQ, GGUF, QLoRA β fit LLMs into the memory you have. The course behind our KV-cache measurements
Free Benchmark Code
The scripts and raw results behind the measurement posts, starting with the KV cache harness. Free with an account
Premium Series
140 deep-dive posts across 21 series, bilingual KO/EN, with production-ready code and notebooks
Latest Posts
View All β
GPT-2 with RoPE Instead of Learned Position Embeddings: 0.122 Lower Loss After 1B Tokens, 5% Slower
Three GPT-2 (124M) runs with rotary position embeddings ended at 3.552 mean validation loss after 1B FineWeb-Edu tokens, against 3.674 for three runs with GPT-2's learned position embeddings. The gap clears this series' pre-registered bar (0.012) by ten times, at GPT-2's learning rate. RoPE ran 5.3% slower on this code.

Same GPT-2, Same Data, Three Seeds: The Final Loss Spread 0.018 Nats
Three GPT-2 (124M) runs on 1B FineWeb-Edu tokens that differed only in the random seed ended at 3.668, 3.669 and 3.686 validation loss (sd 0.010). Seed 0 run again with the same command landed 0.0035 away. With three seeds per side, a change has to beat 0.016 nats before this series calls it real.

Karpathy's βLet's Reproduce GPT-2β in One Read: The Four Parts, Their Numbers, and What Changed Since
A condensed walk through the 4-hour video: building GPT-2 (124M), taking a step from 1,000 ms to 90 ms on one A100, the GPT-3 training settings, and a 10B-token run that beat OpenAI's checkpoint on HellaSwag. Plus one check we ran: nanoGPT's exact GELU moves logits by up to 5 against GPT-2's.

Strata vs llama.cpp on the Same Qwen3.8-Flash-Next File: No Score Difference Found, Faster Within 12 GiB
Same IQ2_XS file in both engines: GSM8K 93.3% vs 93.0%, HumanEval 157/164 each, no vision gap showed up on 100 COCO images. Held to 12 GiB, Strata (draft decoding on) decoded 62.9 tok/s to llama.cpp's 23.4.

llama.cpp KV Cache Quantization: Which -ctk and -ctv to Use, Measured on the Current Build
Keep f16 if the KV cache fits. If not, set -ctk q8_0 -ctv q8_0: it saved 4.2 GiB at 64K but decoded at 56% of f16's speed there.

What Are You Deciding, and Which Mistakes Are Expensive? Chapter 1 of a Book on Building Decision Systems
Free chapter: before training anything, count what each intent can tell you, find the intents likely to be confused, and write down what each mistake costs. The code runs in under a second.