Leaderboard
Models ranked by Taste Elo from same-task duels, with mean rubric scores per dimension, slop density, and duel counts. Each model opens a page showing exactly how it was judged; the protocol is on the Methodology page.
| Rank | Model | Taste Elo | Clarity & structure | Idiom & economy | Signal discipline | Communication | Test taste | Slop density | Duels |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 (Anthropic) | 1606.4 | 3.25 | 2.96 | 3.06 | 3.34 | 3.48 | 16.0% | 27 |
| 2 | Kimi K3 (Moonshot) | 1595.1 | 3.15 | 3.03 | 2.95 | 3.26 | 3.46 | 19.6% | 22 |
| 3 | GLM 5.3 (Z.ai) | 1516.6 | 3.26 | 3.01 | 2.88 | 3.35 | 3.48 | 17.5% | 20 |
| 4 | GPT-6 Astra (OpenAI) | 1465.7 | 3.02 | 2.82 | 2.82 | 3.13 | 3.37 | 17.3% | 30 |
| 5 | Grok 4.7 (xAI) | 1432.1 | 3.08 | 2.94 | 3.29 | 3.23 | 3.29 | 15.2% | 25 |
| 6 | Gemini 3.8 Flash (Google) | 1384.0 | 3.17 | 2.89 | 2.65 | 2.40 | 3.10 | 18.7% | 28 |
Run details
Judgejev-1.13
Swap agreement60.5%
Repeat agreement100.0%
Duels judged76