Learn AI by Building
From your first dataset to production agents โ deep-dive series, hands-on notebooks, and experiments you can rerun yourself.
Paper of the Week
All Issues โPaper of the Week #6 โ The Check Passed. What Did It Check?
This week: a Lean proof that does not match the paper it formalizes, a decision-model table whose test sets are in its training mix, labels that override definitions, memory that wins or loses depending on how much the model reads, NCCL symmetric memory gains that depend on payload size and GPU, a 40-49% token cut that needs lowercase and thinking off, and an AI prescribing pilot whose first phase has two physicians check every prescription.
Premium Series
Our Products
Courses and code built from what we measure here
Video Courses
13 hands-on courses โ quantization, diffusion, RAG, on-device AI, decision models. $199 lifetime bundle, first 3 lectures of every course free
LLM Quantization and Compression Hands-On
GPTQ, AWQ, GGUF, QLoRA โ fit LLMs into the memory you have. The course behind our KV-cache measurements
Free Benchmark Code
The scripts and raw results behind the measurement posts, starting with the KV cache harness. Free with an account
Premium Series
140 deep-dive posts across 21 series, bilingual KO/EN, with production-ready code and notebooks
Latest Posts
View All โ
How Much of an Uncertainty Score Is Just Answer Length? U-Space on Qwen3.5-4B
A new paper finds that answer length alone flags a language model's wrong answers almost as well as many uncertainty scores. We ran the authors' code on Qwen3.5-4B with their items: length alone reached AUROC 65.3 on three benchmarks, from 75.2 on TriviaQA to 48.6 on SuperGPQA. After controlling for length, U-Lens beat MSP (+8.4) and length alone (+4.4 before control), but its edge over mean log-loss (+1.1) was too small to call. On the three benchmarks in the main comparisons up to a third of answers hit the 16,384-token limit and were excluded; on Omni-MATH, left out of the comparisons, 92.5% did.

How Far Can You Shrink a Decision Model? decider-4b from BF16 to Q2_K on llama.cpp
We ran the open decision model decider-4b at five GGUF sizes, from BF16 (10.4 GiB of GPU memory) to Q2_K (4.4 GiB), on 500 TREC questions. Accuracy differences were too small to detect at every size. The probabilities moved: log-loss got slightly better at Q4_K_M and worse at Q8_0 (barely), Q3_K_M and Q2_K. Q2_K changed 13% of the answers and Q3_K_M cut the share of answers above 0.9 from 40% to 23%. Q4_K_M halved the memory; no accuracy difference was detected and its log-loss improved.

Can You Act on Microsoft-Decision-1's Probabilities? We Swapped the Labels on 500 Questions
We sent 500 TREC questions to Microsoft-Decision-1 and Jev 1.13 through OpenRouter, and ran two local decision models, decider-4b and laya, with normal labels and with option names swapped against their definitions. Decision-1 followed the definitions (95.2% against 96.4%, not a difference by our rule) and at a 0.9 threshold handled 80% of the swapped cases on its own, 5 of those 400 wrong. laya followed the labels and was wrong on 110 of the 115 swapped cases it would have handled. The same request sent twice gave slightly different probabilities on Decision-1.

Calling Microsoft-Decision-1 on OpenRouter: Six Things We Hit on Day One
A short field guide from our first day with Microsoft-Decision-1 through OpenRouter: it uses the System One endpoint, not chat completions; choice options must be a map; confidence is not the top probability; the model ID carries a date; requests hit rate limits; and the same question cost us 3.5 to 4.8 times as much on Jev. Working Python included.

Does Codemode Work With Small Local Models? Qwen3.5-9B Got 28 of 30 Writing Code, 19 Calling Tools
Armin Ronacher writes that codemode, where the model writes code that calls tools instead of calling them one at a time, does not yet work with smaller models. We tested it on 30 tasks against a mock issue tracker with Qwen3.5-4B, Qwen3.5-9B and Qwen3.8-27B in llama.cpp. Codemode got more tasks right at every size (22 vs 18, 28 vs 19, 30 vs 26); only the 9B difference passes our test (p = 0.012), and most of it comes from tool-calling replies cut off at our 2,048-token output limit. On tasks both modes got right, codemode's median tokens per task were 28-45% of tool calling's.

Speculative Decoding Speedups Depend on the Task: Prompt Lookup 6.2ร on a File Edit, DFlash 0.7ร on an Essay (Qwen3.5-9B, llama.cpp)
We measured three kinds of speculative decoding in llama.cpp on one Qwen3.5-9B file, alone on an A100: prompt n-gram lookup, the model's own MTP head, and a DFlash draft model. Rewriting a 130-line file ran 6.2ร faster with n-gram lookup and 3.2ร with DFlash. On a 500-word essay, MTP was 4% faster, n-gram lookup slightly slower and DFlash 0.69ร. At temperature 0, n-gram and MTP output matched the plain run token for token; DFlash output diverged on three of four tasks.