Learn AI by Building

From your first dataset to production agents β€” deep-dive series, hands-on notebooks, and experiments you can rerun yourself.

Paper of the Week

All Issues β†’
Issue #4 Β· WeeklySep 30, 2026

Paper of the Week #4 β€” A Memory of Procedures, or a Memory of Examples?

Designer-RSI grows a natural-language skill bank from user traffic and lifts execution success from 72.7% to 99.3% with no weight updates. I built the narrow version on a task with human labels: 40 rules distilled from the model's own mistakes fixed 4 items and broke 5. Retrieving five raw examples fixed 19 and broke none.

Premium Series

Our Products

Courses and code built from what we measure here

Latest Posts

View All β†’
GPT-2 with RoPE Instead of Learned Position Embeddings: 0.122 Lower Loss After 1B Tokens, 5% Slower

GPT-2 with RoPE Instead of Learned Position Embeddings: 0.122 Lower Loss After 1B Tokens, 5% Slower

Three GPT-2 (124M) runs with rotary position embeddings ended at 3.552 mean validation loss after 1B FineWeb-Edu tokens, against 3.674 for three runs with GPT-2's learned position embeddings. The gap clears this series' pre-registered bar (0.012) by ten times, at GPT-2's learning rate. RoPE ran 5.3% slower on this code.

- Models & Algorithms
Read More
Same GPT-2, Same Data, Three Seeds: The Final Loss Spread 0.018 Nats

Same GPT-2, Same Data, Three Seeds: The Final Loss Spread 0.018 Nats

Three GPT-2 (124M) runs on 1B FineWeb-Edu tokens that differed only in the random seed ended at 3.668, 3.669 and 3.686 validation loss (sd 0.010). Seed 0 run again with the same command landed 0.0035 away. With three seeds per side, a change has to beat 0.016 nats before this series calls it real.

- Models & Algorithms
Read More
Karpathy's β€œLet's Reproduce GPT-2” in One Read: The Four Parts, Their Numbers, and What Changed Since

Karpathy's β€œLet's Reproduce GPT-2” in One Read: The Four Parts, Their Numbers, and What Changed Since

A condensed walk through the 4-hour video: building GPT-2 (124M), taking a step from 1,000 ms to 90 ms on one A100, the GPT-3 training settings, and a 10B-token run that beat OpenAI's checkpoint on HellaSwag. Plus one check we ran: nanoGPT's exact GELU moves logits by up to 5 against GPT-2's.

- Models & Algorithms
Read More
Strata vs llama.cpp on the Same Qwen3.8-Flash-Next File: No Score Difference Found, Faster Within 12 GiB

Strata vs llama.cpp on the Same Qwen3.8-Flash-Next File: No Score Difference Found, Faster Within 12 GiB

Same IQ2_XS file in both engines: GSM8K 93.3% vs 93.0%, HumanEval 157/164 each, no vision gap showed up on 100 COCO images. Held to 12 GiB, Strata (draft decoding on) decoded 62.9 tok/s to llama.cpp's 23.4.

- Models & Algorithms
Read More
llama.cpp KV Cache Quantization: Which -ctk and -ctv to Use, Measured on the Current Build

llama.cpp KV Cache Quantization: Which -ctk and -ctv to Use, Measured on the Current Build

Keep f16 if the KV cache fits. If not, set -ctk q8_0 -ctv q8_0: it saved 4.2 GiB at 64K but decoded at 56% of f16's speed there.

- Models & Algorithms
Read More
What Are You Deciding, and Which Mistakes Are Expensive? Chapter 1 of a Book on Building Decision Systems

What Are You Deciding, and Which Mistakes Are Expensive? Chapter 1 of a Book on Building Decision Systems

Free chapter: before training anything, count what each intent can tell you, find the intents likely to be confused, and write down what each mistake costs. The code runs in under a second.

- Models & Algorithms
Read More