Skip to content
K Koda Intelligence
KODA TERMINAL / THE LEADERBOARD 115 RANKED ENTRIES FILED
INSTRUMENT

Leaderboard.

FILED AS

Public benchmark boards, read from each board's own published table and reconciled to one model registry so a lab spelled two ways is counted once. Read the method.

BOARDS
4
RANKED ENTRIES
115
LABS
14, on canonical identity
SCRAPED
13 SEP 2026
CHAMPIONS · Artificial Analysis
#1
INDEX
Claude Fable 5.1
53
Anthropic
$20.00 /1M67.1 tok/s
LISTED AS Claude Fable 5.1 (max with fallback)
SCALE 20 TO 60 INDEXartificialanalysis.ai
#2
INDEX
GPT-6 Astra
53
OpenAI
$20.00 /1M60.1 tok/s
LISTED AS GPT-6 Astra (max)
SCALE 20 TO 60 INDEXartificialanalysis.ai
#3
INDEX
Claude Opus 5
51
Anthropic
$10.00 /1M52.7 tok/s
LISTED AS Claude Opus 5 (max)
SCALE 20 TO 60 INDEXartificialanalysis.ai
// 01

The Consensus Board

10 families on 2 or more boards · percentile 0 to 100
0255075100
Mean
Gap
Claude Fable 5.1
3/4
96.0
12.0
Claude Fable 5
3/4
94.7
12.5
UNANIMOUS
Muse Spark 1.3
3/4
84.3
9.7
Claude Opus 5
3/4
79.8
19.7
Muse Spark 1.2
2/4
69.6
28.8
SPLIT
GPT-6 Astra
3/4
68.3
79.8
UNANIMOUS
Gemini 3.7 Flash
2/4
66.5
5.0
UNANIMOUS
Grok 4.6
2/4
66.5
8.8
SPLIT
DeepSeek V4.1 Flash
2/4
66.0
40.4
SPLIT
GPT-5.6 Sol
3/4
66.0
46.8
Claude Fable 5.13/496.012.0
Claude Fable 53/494.712.5
UNANIMOUSMuse Spark 1.33/484.39.7
Claude Opus 53/479.819.7
Muse Spark 1.22/469.628.8
SPLITGPT-6 Astra3/468.379.8
UNANIMOUSGemini 3.7 Flash2/466.55.0
UNANIMOUSGrok 4.62/466.58.8
SPLITDeepSeek V4.1 Flash2/466.040.4
SPLITGPT-5.6 Sol3/466.046.8
One board percentileMean of those dotsLowest to highestDot colour is the provider colour used on every panel belowTap or focus a dot for its board, percentile and rank

On each board a family earns a percentile from where it sits among that board's canonical families: top of the board scores 100.0, bottom scores 0.0. The mean is those percentiles in the order above. Gap is the distance from a family's lowest board percentile to its highest. Tags come from the ten gaps on this table: the tightest quarter (10.2 points or less) reads UNANIMOUS, the widest quarter (37.5 or more) reads SPLIT. Today that is 3 unanimous, 3 split, and 4 that the boards have not ruled on. Composite koda-leaderboard-composite-v1.

#ModelBoardsThe Spread0 = board floor, 100 = board topMeanGappointsLMArenaLiveBenchSWE-Bench VerifiedArtificial Analysis
1
Claude Fable 5.1
3/496.012.0#4#1·#1
2
Claude Fable 5
3/494.712.5#1#2·#4
3
UNANIMOUS
Muse Spark 1.3
3/484.39.7#7#4·#5
4
Claude Opus 5
3/479.819.7#10#8·#3
5
Muse Spark 1.2
2/469.628.8#5#14··
6
SPLIT
GPT-6 Astra
3/468.379.8#25#3·#2
7
UNANIMOUS
Gemini 3.7 Flash
2/466.55.0#12#10··
8
UNANIMOUS
Grok 4.6
2/466.58.8·#12·#8
9
SPLIT
DeepSeek V4.1 Flash
2/466.040.4·#5·#14
10
SPLIT
GPT-5.6 Sol
3/466.046.8#20#6·#6
// 02

The Instruments

4 boards · read 13 SEP 2026 · 0 stale

Every board read straight from source. The champion board leads at full width with price and throughput on screen; the other three sit beneath it. Each panel anchors its bars to a band computed from the rows it actually prints, so bar length is distance across that band, never from zero. A score column carries as many decimals as its own board publishes, capped at one, which is why an index ranked in whole points prints whole points beside a score ranked in tenths: no panel invents a resolution it was not given.

Artificial AnalysisIntelligenceREAD TODAYSCRAPED SEP 13, 2026 22:17 UTCartificialanalysis.ai
SCALE 20 TO 60 INDEX · BAND COMPUTED OVER THE 20 ROWS SHOWN · 20 OF 25 READ
1Claude Fable 5.153$20.00 /1M67.1 tok/s
2GPT-6 Astra53$20.00 /1M60.1 tok/s
3Claude Opus 551$10.00 /1M52.7 tok/s
4Claude Fable 550$20.00 /1M64.9 tok/s
5Muse Spark 1.348$2.00 /1M236.7 tok/s
6GPT-5.6 Sol47$8.00 /1M57.9 tok/s
7GLM-5.3245$2.15 /1M72.0 tok/s
8Grok 4.6144$3.00 /1M60.0 tok/s
9Kimi K3144$6.00 /1M36.8 tok/s
10GPT-5.6 Terra142$4.50 /1M113.7 tok/s
11GLM-5.3-Flash242$0.24 /1M108.7 tok/s
12Gemini 3.8 Flash241$1.50 /1M277.5 tok/s
13Qwen3.8 2.4T A95B140$3.00 /1M40.0 tok/s
14DeepSeek V4.1 FlashNEW40$0.52 /1M227.8 tok/s
15GPT-5.6 Luna138$0.45 /1M119.7 tok/s
16DeepSeek V4 Pro 0813136$1.98 /1M77.9 tok/s
17Qwen3.8 27B134$1.12 /1M42.7 tok/s
18K2 Horizon 375B A23B131no pricenot measured
19MiniMax-M3130$0.52 /1M109.3 tok/s
20Inkling126$1.76 /1M89.3 tok/s
LMArenaGeneralREAD TODAYSCRAPED SEP 13, 2026arena.ai
SCALE 1490 TO 1510 ELO · BAND COMPUTED OVER THE 8 ROWS SHOWN · 8 OF 30 READ
1Claude Fable 51506
2Claude Opus 4.6HIGH1505
3Claude Opus 4.7HIGH11502
4Claude Fable 5.111501
5Muse Spark 1.21499
6Claude Opus 4.6STD1497
7Muse Spark 1.3NEW1496
8Claude Opus 4.7STD11494
LiveBenchReasoningREAD TODAYSCRAPED SEP 13, 2026livebench.ai
SCALE 80 TO 84 SCORE · BAND COMPUTED OVER THE 8 ROWS SHOWN · 8 OF 30 READ
1Claude Fable 5.183.4
2Claude Fable 583.0
3GPT-6 Astra82.2
4Muse Spark 1.381.6
5DeepSeek V4.1 FlashNEW81.1
6GPT-5.6 Sol181.0
7GPT-5.5180.2
8Claude Opus 5180.1
SWE-Bench VerifiedCodingREAD TODAYSCRAPED SEP 13, 2026swebench.com
SCALE 72 TO 77 SCORE · BAND COMPUTED OVER THE 8 ROWS SHOWN · 8 OF 30 READ
1Claude 4.5 OpusHIGH76.8
2Gemini 3 Flash75.8
3MiniMax M2.575.8
4Claude Opus 4.675.6
5Claude 4.5 OpusMED74.4
6Gemini 3 Pro Preview74.2
7GLM 572.8
8GPT 5.272.8
// 03

Who Owns The Board

80 slots · top 20 per board

Share of every top-20 canonical family slot across all 4 boards, by provider. One bar, 80 slots, the whole field at a glance. Percentages read as whole numbers on the bar and to one decimal on hover. 2 of those slots belong to labs with no validated palette colour and share the neutral segment: MBZUAI (1), Unknown (1).

Anthropic 25% 20 slotsOpenAI 20% 16 slotsGoogle 14% 11 slots8 more labs hold the remaining 33 slots, the largest 8.8 percent
// 04

Head to Head

percentile, board by board

Both models are lined up on every board they share. The bars are board percentiles, so two different scales can sit on one axis, and the delta is the distance between them: a measurement, not a verdict. Each side wears its own lab colour, or ink where both come from the same lab, and the exact score sits on hover. Boards where one is missing are skipped, not faked.

Delta
// 05

The Colour Contract

one role per hex
First place
F59E0B

The one chromatic gesture on the podium and in the board cells. Where a lab wears the same hue as the medal, its dot renders neutral across the podium and the lab is named in words.

Second and third
#2#3

Ink, not hue. Three drawings on one 24 grid at one stroke weight, so the places separate without colour.

Labs
AnthropicOpenAIGoogle

One validated hue per lab, on every dot, tick, bar and segment. 2 of the 80 slots belong to labs with no validated hue and render neutral rather than borrowing another lab's colour.

Chrome
// 01Generalarena.ai

Ink and hairline. Eyebrows, category labels, source links, focus rings and row hover carry no hue of their own, so a colour inside this instrument always belongs to a lab.

Readings
STALE 9D54.22NEW

Neutral outline, ink text, no fill. Age, head to head delta and rank movement are measurements, not verdicts. Nothing is past the 8 day window today.

Agreement
UNANIMOUSSPLIT

Typographic, no hue. Ink for agreement, muted for disagreement, leading the row.

// 06

How This Is Counted

koda-leaderboard-composite-v1

The stamp cannot precede its rows

The page prints DATA AS OF SEP 13, 2026 22:17 UTC, the newest scrape on it. The composite stamp is at or after every row it covers, so the Dateline square is filled.

Identity is a family, not a string

Each board contributes one row per family, its best ranked configuration. The raw board string stays beside it as LISTED AS, and where one board ranks two settings of the same product the rows carry the effort marker the registry recorded rather than repeating the name.

Age is recomputed in your browser

Every board carries its scrape stamp, so a page that stops rebuilding still ages on screen. The server value is the no-script fallback. Stale after 8 days.

Different boards measure different things

The consensus view averages family percentiles, not scores, and each panel measures its bars across a band computed from its own rows. Provider colours come from one validated palette; a lab with no entry renders neutral rather than borrowing another lab's colour. Sources: arena.ai, livebench.ai, swebench.com, artificialanalysis.ai.

LMArenaarena.aiSCRAPED SEP 13, 2026READ TODAY

Human preference rankings from blind A/B comparisons (7M+ votes)

LiveBenchlivebench.aiSCRAPED SEP 13, 2026READ TODAY

Contamination-free benchmark with monthly-refreshed questions

SWE-Bench Verifiedswebench.comSCRAPED SEP 13, 2026READ TODAY

Real-world software engineering task completion (Verified, 500 tasks)

Artificial Analysisartificialanalysis.aiSCRAPED SEP 13, 2026READ TODAY

Artificial Analysis Intelligence Index v4.3: aggregate of 10 evaluations

Get the morning Signal

171 editions so far, one a day. Unsubscribe anytime.