AHD Artificial Human Design

AHD · Evals

Every run.

One row per published surface. Most rows are a single dated run. The weekly row is a series covering eight of them, with its own register linking each dated report. Every claim on this site traces back to one of those reports.

Cell, raw, compiled, n, tells: the glossary defines every term these rows use.

  1. Eight runs · Same brief

    Weekly series · CF OSS n=30 Token swiss-editorial n=30 2400 samples

    The weekly programme, read as a series rather than eight bulletins. gpt-oss and mistral reduce every run and by a consistent amount. gemma reduces every run by an amount that swings twenty points. llama-4-scout sits against zero. qwen3 changes sign six times in eight runs, so no claim about its direction survives the series. Running weekly is what separates a cell pinned near zero from a cell oscillating around it.

  2. Eleven models · Same brief · Different token

    Cross-provider n=30 · post-digital-green Token post-digital-green n=30 660 samples

    The different-token-same-brief triangulation queued by the 22 April report. Eight of eleven cells regress under the compiled prompt. The compiler is not at fault: it transmits the token faithfully. The regressions come from lint rules that assumed the editorial defaults this token rejects, so they penalised output that followed the token. gpt-5.5 lands as the cleanest raw frontier baseline measured to date at 1.03 tells per page. Token-aware linting is the next engineering step.

  3. Ten models · One brief · Thirty samples each

    Cross-provider n=30 Token swiss-editorial n=30 600 samples

    Ten models, n=30 per cell, 600 samples. Eight of ten cells showed positive reductions under the compiled prompt, one flat, one regression. Best: gpt-oss-120b at 78.1% fewer tells. Three frontier cells via subscription CLIs (Claude Code, Codex, Gemini CLI), seven OSS via Cloudflare Workers AI. Wilson interval tightens from roughly ±35% at n=5 to roughly ±18% at n=30.

  4. Seven models across four providers

    Cross-provider n=5 Token swiss-editorial n=5 70 samples

    Four positive reductions, one inconclusive, two regressions. Llama 3.3's regression reproduces across Cloudflare and Hugging Face, turning a single-cell finding into a cross-provider result.

  5. Five models · n=5 · Zero errors

    Five-model n=5 Token swiss-editorial n=5 50 samples

    Claude Opus (Anthropic API) plus four OSS models on Cloudflare Workers AI. Claude dropped to zero tells compiled; Llama 3.3 70B regressed. Full per-model and per-tell breakdown with every attempted-vs-scored count published.

What's in each report

A dated run page carries a summary table of raw versus compiled tells per cell, the attempted / extracted / scored column triple for every cell, a per-tell frequency table showing which rules fired in which conditions, the exact brief and compiled-prompt bytes used, plus the run manifest with canonical model identifiers and serving paths. The weekly series page carries the reduction for every cell in every run and a register linking each dated report in the framework repository, where those manifests and per-tell tables live. It does not restate them, because a per-tell table averaged across eight runs would hide the variance the series exists to show. The rows are headlines. The evidence is in the reports.

What a run does not carry

A run is not a leaderboard. A run is evidence for a specific brief under a specific token against a specific set of models. Different briefs and different tokens produce different rankings. If you generalise from one row to "model X is better than model Y at design", that claim is yours. The run does not support it. The frame is methodology. Every cell names its serving path because a model served by two hosts is two targets.

Contribute a run

If you have budget, keys and a model we haven't measured, you can add to this record. The short version of what a submittable run looks like (full manifest, provider request-IDs, n ≥ 3, no post-processing, negative results reported) lives on the submission page. The full protocol lives in the framework repo's CONTRIBUTING.md.

Adjacent: methodology, positioning, the taxonomy, contribute a run.