4 boards. 97 models ranked across 15 labs. One consensus verdict, verified weekly from the original sources.
The cross-board verdict. For each board a model family earns a family percentile (1 minus its canonical position), and the composite is the mean percentile across every board it appears on. Shown for models ranked on two or more boards. koda-leaderboard-composite-v1 · as of AUG 02, 2026 01:58 UTC.
| # | Model | Composite | LMArena | LiveBench | SWE-Bench Verified | Artificial Analysis |
|---|---|---|---|---|---|---|
| 1 | claude-fable-5 | 98.6 | #1 | #1 | · | #2 |
| 2 | Claude Opus 5 (max) | 92.2 | #6 | #4 | · | #1 |
| 3 | claude-opus-4-7-thinking | 81.9 | #3 | #9 | · | · |
| 4 | GPT-5.6 Sol Max Effort | 81.5 | #14 | #2 | · | #3 |
| 5 | GPT-5.6 Terra (max) | 81.0 | · | #7 | · | #5 |
| 6 | Kimi K3 (max) | 79.5 | #12 | #5 | · | #4 |
| 7 | claude-opus-4-6-thinking | 73.0 | #2 | #16 | #4 | · |
| 8 | Gemini 3.1 Pro Preview High | 72.7 | #11 | #8 | · | · |
| 9 | GPT-5.5 Thinking xHigh Effort | 70.5 | #16 | #3 | · | · |
| 10 | Grok 4.5 (high) | 70.2 | · | #12 | · | #6 |
Every board, read straight from source. Bars scale to each board’s top score, so length is proportional and the exact figure sits in ink at the end.
Share of every top-20 canonical family slot across all 4 boards, by provider. One bar, 73 slots, the whole field at a glance.
Pick any two canonical model families. We line them up on every board they share and show the exact configuration used. Boards where one is missing are skipped, not faked.
Composite koda-leaderboard-composite-v1 · as of AUG 02, 2026 01:58 UTC. Every figure is pulled from the source below and verified before it ships.
Different boards measure different things; the consensus view averages canonical-family percentiles, not scores. Each family contributes once per board using its best-ranked configuration.