Leaderboard

Models ranked by Taste Elo from same-task duels, with mean rubric scores per dimension, slop density, and duel counts. Each model opens a page showing exactly how it was judged; the protocol is on the Methodology page.

Models ranked by Taste Elo, with per-dimension rubric means, slop density, and duel counts
RankModelTaste EloClarity & structureIdiom & economySignal disciplineCommunicationTest tasteSlop densityDuels
1 Claude Fable 5.1 (Anthropic) 1606.4 3.252.963.063.343.48 16.0% 27
2 Kimi K3 (Moonshot) 1595.1 3.153.032.953.263.46 19.6% 22
3 GLM 5.3 (Z.ai) 1516.6 3.263.012.883.353.48 17.5% 20
4 GPT-6 Astra (OpenAI) 1465.7 3.022.822.823.133.37 17.3% 30
5 Grok 4.7 (xAI) 1432.1 3.082.943.293.233.29 15.2% 25
6 Gemini 3.8 Flash (Google) 1384.0 3.172.892.652.403.10 18.7% 28

Run details

Judgejev-1.13
Swap agreement60.5%
Repeat agreement100.0%
Duels judged76

Swap and repeat agreement are judge reliability — how reproducible the judge is — not agreement with human raters. Repeat sample: 7 duels.

Sealed tier, in aggregate

4 tasks · 20 items · 43 duels. Sealed briefs, code, and prompts never enter this bundle; only these counts and the published score distributions ship publicly.

Generated 2026-09-22T14:54:35.963Z · suite 561466d2340d44bace6b92f320363ac45b89d787