AHD · Evals
Every run.
One row per published surface. Most rows are a single dated run. The weekly row is a series covering eight of them, with its own register linking each dated report. Every claim on this site traces back to one of those reports.
Cell, raw, compiled, n, tells: the glossary defines every term these rows use.
-
Eight runs · Same brief
The weekly programme, read as a series rather than eight bulletins. gpt-oss and mistral reduce every run and by a consistent amount. gemma reduces every run by an amount that swings twenty points. llama-4-scout sits against zero. qwen3 changes sign six times in eight runs, so no claim about its direction survives the series. Running weekly is what separates a cell pinned near zero from a cell oscillating around it.
-
Eleven models · Same brief · Different token
The different-token-same-brief triangulation queued by the 22 April report. Eight of eleven cells regress under the compiled prompt. The compiler is not at fault: it transmits the token faithfully. The regressions come from lint rules that assumed the editorial defaults this token rejects, so they penalised output that followed the token. gpt-5.5 lands as the cleanest raw frontier baseline measured to date at 1.03 tells per page. Token-aware linting is the next engineering step.
-
Ten models · One brief · Thirty samples each
Ten models, n=30 per cell, 600 samples. Eight of ten cells showed positive reductions under the compiled prompt, one flat, one regression. Best: gpt-oss-120b at 78.1% fewer tells. Three frontier cells via subscription CLIs (Claude Code, Codex, Gemini CLI), seven OSS via Cloudflare Workers AI. Wilson interval tightens from roughly ±35% at n=5 to roughly ±18% at n=30.
-
Seven models across four providers
Four positive reductions, one inconclusive, two regressions. Llama 3.3's regression reproduces across Cloudflare and Hugging Face, turning a single-cell finding into a cross-provider result.
-
Five models · n=5 · Zero errors
Claude Opus (Anthropic API) plus four OSS models on Cloudflare Workers AI. Claude dropped to zero tells compiled; Llama 3.3 70B regressed. Full per-model and per-tell breakdown with every attempted-vs-scored count published.
What's in each report
A dated run page carries a summary table of raw versus compiled tells per cell, the attempted / extracted / scored column triple for every cell, a per-tell frequency table showing which rules fired in which conditions, the exact brief and compiled-prompt bytes used, plus the run manifest with canonical model identifiers and serving paths. The weekly series page carries the reduction for every cell in every run and a register linking each dated report in the framework repository, where those manifests and per-tell tables live. It does not restate them, because a per-tell table averaged across eight runs would hide the variance the series exists to show. The rows are headlines. The evidence is in the reports.
What a run does not carry
A run is not a leaderboard. A run is evidence for a specific brief under a specific token against a specific set of models. Different briefs and different tokens produce different rankings. If you generalise from one row to "model X is better than model Y at design", that claim is yours. The run does not support it. The frame is methodology. Every cell names its serving path because a model served by two hosts is two targets.
Contribute a run
If you have budget, keys and a model we haven't measured,
you can add to this record. The short version of what a
submittable run looks like (full manifest, provider
request-IDs, n ≥ 3, no post-processing, negative results
reported) lives on the
submission page. The full
protocol lives in the framework repo's
CONTRIBUTING.md.
Adjacent: methodology, positioning, the taxonomy, contribute a run.