Learn AI by Building
From your first dataset to production agents โ deep-dive series, hands-on notebooks, and experiments you can rerun yourself.
Tutorials
View All โLLM Agent Cookbook
Build AI agents from scratch โ ReAct, Tool Use, Multi-Agent orchestration
ML Cookbook
Master machine learning algorithms with hands-on Jupyter projects
Data Analysis Cookbook
SQL, Pandas, Statistics โ everything for data-driven decisions
Ontology & KG Cookbook
RDF, OWL, Neo4j, and GraphRAG for knowledge-powered AI
Paper of the Week
All Issues โPaper of the Week #4 โ A Memory of Procedures, or a Memory of Examples?
Designer-RSI grows a natural-language skill bank from user traffic and lifts execution success from 72.7% to 99.3% with no weight updates. I built the narrow version on a task with human labels: 40 rules distilled from the model's own mistakes fixed 4 items and broke 5. Retrieving five raw examples fixed 19 and broke none.
Premium Series
Our Products
Courses and starter kits built from what we measure here
Video Courses
13 hands-on courses โ quantization, diffusion, RAG, on-device AI, decision models. $199 lifetime bundle, first 3 lectures of every course free
LLM Quantization and Compression Hands-On
GPTQ, AWQ, GGUF, QLoRA โ fit LLMs into the memory you have. The course behind our KV-cache measurements
Starter Kits
Solution notebooks for the free cookbooks โ LLM Agent, Data Analyst, ML, Ontology & KG. One-time purchase
Premium Series
140 deep-dive posts across 21 series, bilingual KO/EN, with production-ready code and notebooks
Starter Kits
View All โPractice notebooks, interview questions, and project solutions โ ready to download.
Browse Starter KitsLatest Posts
View All โ
Paper of the Week #4 โ A Memory of Procedures, or a Memory of Examples?
Designer-RSI grows a natural-language skill bank from user traffic and lifts execution success from 72.7% to 99.3% with no weight updates. I built the narrow version on a task with human labels: 40 rules distilled from the model's own mistakes fixed 4 items and broke 5. Retrieving five raw examples fixed 19 and broke none.

Jeff vs Jev on the Same Questions: Overall Scores and a 26-Option Limit (Fixed in v1.1)
On Jeff's own 4,599 questions, Jev scored 85.7 and Jeff-2B 83.0. Jeff v1.0 never picked an option past the 26th in my tests; v1.1 fixes that, remeasured.

Kev vs Jev: Kev 0.8B to 9B Benchmarked for Accuracy and Calibration
Kev (0.8B, 4B, 9B) and Jev 1.13 on the same 1,629 labelled messages (2,329 requests). On the three datasets Kev trained on, Kev-9B was as accurate or more (TREC 93.8% vs 89.0%). On three it never saw, Jev was ahead on two (CLINC150 68.5% vs 62.0%, MASSIVE 82.9% vs 76.0%). Jev's probabilities ran high and come rounded to two decimals, so on BANKING77 no threshold reached 95% accuracy.

How to Run Qwen Locally: Which File Fits Your GPU (8, 12, 16 or 24 GB), Measured
Which Qwen GGUF fits an 8, 12, 16 or 24 GB GPU, with measured memory: 9B Q4_K_M 5.8 GiB, 27B Q3_K_XL 13.1 GiB, 35B-A3B 7.3 GiB with 30 layers' experts on the CPU.

How Many Labels Is a Decision Model Worth? Kev on Three Tasks It Never Saw
On CLINC150, MASSIVE and financial-news tweets, none of which Kev trained on, Kev-9B matched a small classifier trained on roughly 2 to 20 labelled examples per class. Its probabilities ran too low: stated confidence sat 10โ17 points under its accuracy (calibration error 11โ17%, against 2โ9% on familiar data), so a 95%-accuracy threshold passed only 23โ58% of messages. Out-of-scope questions got low probabilities: 5 of 100 passed that threshold. The 0.8B model answered 'Financials' to 152 of 200 tweets.

Kev vs MoJev: Two More Open Jev Alternatives, Accuracy and Confidence Measured
Kev (0.8B to 9B) and MoJev (0.85B) both speak Jev's API. On BANKING77, TREC and AG News with human labels, Kev's probabilities were usable: at 95% accuracy it could auto-accept 53โ61% of BANKING77, where laya managed 0%. But Kev was trained on these three datasets, and a small classifier trained on the same data still beat it on BANKING77. MoJev reports 0.79% calibration error; here it was 6โ22%, with probabilities too low.