K Koda Intelligence
DEEP DIVE DEEP DIVE № 221 · 05 October 2026DOCKODA-20261005-220E39C12200sha-256 of date + article + 24 checked + 11 computed

A 3% gap turns your AI model into a swappable supplier

DeepSeek's smallest new model, V4.1 Flash, scored 81.1 25 on LiveBench against 83.4 for Anthropic's best. Bloomberg Intelligence puts the China-U.S. benchmark gap at about 3% 14, down from about 15% early in 2026. Input costs $0.15 05 per million tokens off-peak. We think builders should price every model by cost per successful task and re-bid the supplier every quarter.

6 MIN READ · BY THE KODA EDITORIAL TEAM · STRATEGY · MODEL ROUTING
V4.1 FLASH LIVEBENCH81.1VERIFIED CLAIM 01LiveBench
TOP U.S. MODEL83.4REPORTED CLAIM 02Anthropic
CHINA-U.S. GAP3%VERIFIED CLAIM 12↓ 12 PTS FROM 15%
V4.1 FLASH LIVEBENCH81.1LiveBench TOP U.S. MODEL83.4Anthropic CHINA-U.S. GAP3%↓ 12 PTS FROM 15% FLASH INPUT PRICE$0.15PER M TOKENS OFF-PEAK FLASH OUTPUT PRICE$0.60PER M TOKENS OFF-PEAK NOISE BAND2.7 PTS1,436 TEST ITEMS GPT-4-CLASS PRICE$0.40↓ 98.7% FROM $30 INFERENCE MARKET 2030$255B↑ 2.5× FROM $103B

DeepSeek's smallest new model scored 81.1 01 on LiveBench in September. Anthropic's best model scored 83.4 02. The gap is 2.3 25 points, about 2.8%.

The model is also cheap. DeepSeek's rate card lists DeepSeek-V4.1-Flash at $0.15 05 per million cache-miss input tokens ($0.003 per million on cache hits) and $0.60 per million output tokens off-peak. A job of 1 billion 26 input tokens and 100 million output tokens costs about $210. At peak rates it costs $420 27.

We should own the weak spot up front. Nobody can say for sure that a 2.3 25-point gap means anything. It may be noise.

If capability keeps converging at that pace, the model you pick today is a supplier you will re-bid every quarter.

Cost Per Successful Task Is the Unit

A leaderboard tells you how a model did on someone else's exam. A builder pays for something narrower: outputs that pass and ship. Several 2026 analyses of the DeepSeek release settle on the same yardstick for that: cost per successful task. Take everything you spent to get answers and divide by the answers you accepted. That spend counts retries and human review too.

FLASH PRICE CHECK · OCTOBER 2026DEEPSEEK RATE CARD · The RundownBASE: 24 CHECKED + 11 COMPUTED, 4 SHOWN

What a 1-billion-token job costs on V4.1 Flash

Input tokens DeepSeek · per million, cache miss, off-peak REPORTED CLAIM 05
$0.15
Output tokens DeepSeek · per million, off-peak VERIFIED CLAIM 06
$0.60
Sample job, off-peak 1B input plus 100M output tokens COMPUTED CLAIM 26
$210
Sample job, peak Weekdays 01:00-04:00 and 06:00-10:00 UTC COMPUTED CLAIM 27
$420

Every fact in this story sorts into one side of that fraction. The top half is spend. DeepSeek's off-peak input price is $0.15 05 per million tokens, with cache hits at $0.003. Peak hours double every rate, according to DeepSeek and The Rundown. Those hours run 01 07:00 to 04:00 and 06:00 to 10:00 UTC on weekdays.

DeepSeek's launch post says V4.1 Flash's KV cache needs a quarter of the high-bandwidth memory and an eighth of the SSD storage of the prior generation. The KV cache is the model's working memory of the conversation so far, and agents reread it constantly. DeepSeek says cache charges often make up a large share of agent costs.

The bottom half is accepted work. That is where cheap models can lose. Errors compound across agent steps. A 20 28-step agent that succeeds 99% of the time per step finishes about 82% of runs. At 98% 29 per step it finishes about 67%.

So model choice behaves like a price-performance setting. When two models land the same share of accepted outputs on your tasks, the cheaper one wins the fraction. When their pass rates differ, the leaderboard gap stops mattering, and your own numbers decide.

Why a 2.3-Point Gap Is Thin Evidence

The Flash label is the first sign that something has shifted. DeepSeek calls V4.1 Flash the smallest model in its new architecture family. It is a 552 21-billion-parameter mixture-of-experts model that wakes only about 8 billion 22 parameters to read input and 16 billion to write output. A mixture-of-experts model sends each token to a few specialist sub-networks, so most of the model sits idle on any given step.

Still, the tie on paper is shakier than the headline suggests. One 2026 benchmark comparison estimated that with about 1,436 13 test items, gaps under roughly 2.7 points can be treated as noise. The DeepSeek-Anthropic gap is 2.3 25. It is unclear whether the ranking between the two would survive a rerun.

DeepSeek's own numbers show the same pattern. At maximum effort, V4.1 Flash scored 74.2 15 on DeepSWE v1.1, against 74.0 16 for Claude Opus 5 and 73.0 17 for GPT-5.6 Sol, The Rundown reported. Those margins are 0.2 30 and 1.2 points. DeepSeek's changelog also lists 90.6 18 on Terminal-Bench 2.1 and only 30.0 19 on Terminal-Bench 3.0.

Shape matters more than the average. V4.1 Flash Max Effort's LiveBench category scores run from a low of 70.0 20 to a high of 93.3. Your product lives in one or two of those categories. The aggregate only earns the model a trial on your tasks.

The sharpest case study came from DeepSeek itself, in the same week. Its September 9 launch notice said every deepseek-v4-pro request would route to V4.1 Flash from 04 09:00 UTC on September 14. Then its changelog reversed course "in response to user demand," keeping V4 Pro with billing unchanged. Yotta Labs notes there is still no date for a V4.1 Pro.

Swapping, in other words, runs in both directions. Providers change what sits behind a model name, sometimes within a week. A team that pins versions and reruns its tests notices. A team that hard-codes one provider finds out from its customers.

Cash out the door is the honest measure here. A token price is a list price. Retry calls and reviewer hours are what land on the invoice. A cheap model that pushes more cases to humans can cost more per accepted output than a premium one.

Jurisdiction is the other filter. DeepSeek is a Chinese company, and for buyers bound by procurement rules or data-residency laws, that outweighs any token price. The MIT-licensed weights on Hugging Face allow self-hosting. Yotta Labs puts the checkpoint at about 510 GB 23, three times V4 Flash, which changes the hosting math.

The evidence supports one decision today. Build so the model is a setting, and send high-volume, checkable work to whichever model clears your tests cheapest. Three things would change the call: your tests failing on high-cost cases, a U.S. release reopening the gap past the noise band, or DeepSeek raising prices. LiveBench refreshes its questions about every six months, so the next snapshot could look different.

Why the leaderboard tie needs your own tests

INSIDE THE NOISE
2.3 25

The LiveBench gap sits under the noise band.

One 2026 comparison estimated that with about 1,436 13 test items, gaps under roughly 2.7 points can be noise. The DeepSeek-Anthropic gap is 2.3 points, so the ranking may not survive a rerun.

COMPOUNDING ERRORS
67% 29

Small per-step misses sink long agent runs.

A 20 28-step agent at 99% per step finishes about 82% of runs. At 98% per step it finishes about 67%, and every failed run adds retries and reviewer hours to the bill.

NAMES CAN SHIFT
SEP 14

A model name can change meaning within a week.

DeepSeek planned to route every deepseek-v4-pro request to V4.1 Flash from September 14, then reversed course in response to user demand. Teams that pin versions and rerun tests catch these moves early.

2031. The convergence has a slope, and the slope is what matters strategically. Bloomberg Intelligence measured the gap between Chinese and U.S. models' benchmark scores at about 15% 14 early in 2026 and about 3% by September. That is 12 31 points in roughly eight months. At that pace, the top tier sits inside the 2.7 13-point noise band within months, and a leaderboard can no longer tell the models apart.

Price has a longer record. One 2026 analysis tracks GPT-4-class capability from about $30 32 per million tokens in March 2023 to about $0.40. That is a 98.7% drop, or 75 times cheaper in about three years. Repeat it once more by 2031 and the same capability costs about half a cent per million tokens.

Cheap tokens have not shrunk the bill. The same analysis sizes the inference market at $103 billion 34 in 2025 and projects $255 billion by 2030, about 2.5 times larger. If unit prices fall while total spend rises, volume is growing faster than price is falling. More agents and longer contexts absorb the savings.

That math sets up an asymmetric bet. A model gateway is a one-time cost, measured in days of engineering. Lock-in is a cost that grows with every token you send, and the token count is heading up. If convergence continues, each cheaper model becomes a config change. If it reverses, you have spent a few days.

We think the durable asset by 2031 is your own evaluation set. A 2026 analysis argues model capability is becoming less defensible as a standalone moat as prices fall and gaps narrow. Every failed case you log makes the next switch faster and safer, and that compounds. The model becomes a supplier you re-bid. The tests and customer data stay yours.

Wire a Two-Model Router With Evals

Step one: build the gateway. A gateway here means one function that every model call passes through, reading the model name from a config file. Keep all provider-specific code inside it. Point it at your current model first so nothing changes for users.

Step two: build a golden set. A golden set is a fixed list of real inputs with known good answers. Pull 200 35 requests from last month's logs, including the messy ones that broke things before. Write a pass-or-fail check for each one, such as valid JSON with correct field values.

Step three: run both models on the set. Call your current model and deepseek-flash, with versions pinned. Log input tokens, output tokens, latency and the pass-or-fail result for every case. Some cases will fail on both models, and those failures are the most useful rows in the file.

Step four: price the results. Multiply tokens by each provider's rate, then add reviewer minutes at your hourly cost for every failed case. Divide by passed cases. That number is your cost per successful task, and it drives the routing decision.

Step five: route by task type. Send a task type to deepseek-flash only where its pass rate matches your current model and your data rules allow a Chinese provider. Schedule its batch jobs outside 01 08:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, when DeepSeek charges half. Keep high-cost cases on whichever model wins them.

Step six: rerun the set every week. A scheduled job that replays 200 cases takes minutes. DeepSeek's September plan to reroute deepseek-v4-pro showed a model name can change meaning within a week. The weekly rerun is how you catch the next change before a customer does.

DOJO · BUILD THIS WEEKEND

Wire a two-model router and price it by accepted output

  1. Build one gateway function. Route every model call through a single function that reads the model name from a config file, and point it at your current model first so users see no change.
  2. Assemble a 200-case golden set. Pull 200 real requests from last month's logs, including the messy ones, and write a pass-or-fail check for each, such as valid JSON with correct field values.
  3. Price and route by task type. Run both models with pinned versions, add reviewer minutes for each failure, and divide total spend by passed cases. Send a task type to deepseek-flash only where pass rates match and your data rules allow a Chinese provider, then rerun the set weekly.
Practice: Design an AI-Assisted Workflow
THE BOTTOM LINE

Own the tests and re-bid the model every quarter.

V4.1 Flash lands within 2.3 25 points of the top U.S. model at $0.15 05 per million input tokens, and the China-U.S. gap fell from about 15% 14 to about 3% 12 in roughly eight months. A gateway costs a few days of engineering, while lock-in grows with every token you send. Your golden set and failure log compound with each switch. We would build so the model is a setting and let cost per successful task decide.

LISTEN · AUDIO BRIEFINGThe conversation · ~19 min
WATCH · VISUAL NARRATIVEAnimated breakdown · ~8 min
PLAY · YOUTUBE
EDITORIAL RECEIPTKODA-20261005-220E39C12200
As of05 October 2026MethodClaim extraction, dated-evidence review, and temporal consistency gate.CorrectionsContact the Koda desk
EVIDENCE24 CHECKED + 11 COMPUTED · 10 VERIFIED · 13 REPORTED · 1 FAILED
10 verified13 reported1 failed11 computed
  1. 01DeepSeek V4.1 Flash, DeepSeek's smallest new model, scored 81.1 on LiveBench in September 2026VERIFIEDTRUEBENCHMARKapi-docs.deepseek.com
  2. 02Anthropic's best model scored 83.4 on LiveBench in September 2026REPORTEDMOSTLY TRUEBENCHMARKstackfutures.com
  3. 03Bloomberg Intelligence ranked DeepSeek V4.1 Flash sixth in the worldREPORTEDMOSTLY TRUEATTRIBUTIONbloomberg.com
  4. 04Claim removed during the check; its text is not republished.REPORTEDMIXEDATTRIBUTIONCUT FROM COPYbusinesstimes.com.sg
  5. 05DeepSeek's rate card lists DeepSeek-V4.1-Flash at $0.15 per million cache-miss input tokens ($0.003 per million on cache hits) and $0.60 per million output tokens off-peak.REPORTEDMOSTLY TRUEPRICECORRECTED IN COPYdeepseek.ai
  6. 06DeepSeek's rate card lists V4.1 Flash at $0.60 per million output tokens off-peakVERIFIEDTRUEPRICEdeepseek.ai
  7. 07DeepSeek's peak hours, 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, double every rate, according to DeepSeek and The RundownREPORTEDMOSTLY TRUEPRICEtherundown.ai
  8. 08DeepSeek charges half its peak rates outside 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdaysREPORTEDMOSTLY TRUEPRICEapi-docs.deepseek.com
  9. 09DeepSeek's September 9, 2026 launch notice said every deepseek-v4-pro request would route to V4.1 Flash from 04:00 UTC on September 14VERIFIEDTRUEHISTORYdeepseek.com
  10. 10Bloomberg Intelligence put the China-U.S. AI model capability gap near 15% early in 2026REPORTEDMIXEDSTATbloomberg.com
  11. 11Claim removed during the check; its text is not republished.FAILEDMOSTLY FALSESTATCUT FROM COPYbloomberg.com
  12. 12Bloomberg Intelligence put the China-U.S. AI model capability gap at roughly 3% after the DeepSeek V4.1 Flash releaseVERIFIEDTRUESTATbloomberg.com
  13. 13One 2026 benchmark comparison estimated that with about 1,436 test items, gaps under roughly 2.7 points can be treated as noise.REPORTEDMOSTLY TRUESTATCORRECTED IN COPYthemodelgap.com
  14. 14Bloomberg Intelligence measured the gap between Chinese and U.S. models' benchmark scores at about 15% early in 2026 and about 3% by September.REPORTEDMOSTLY TRUESTATCORRECTED IN COPYbloomberg.com
  15. 15At maximum effort, DeepSeek V4.1 Flash scored 74.2 on DeepSWE v1.1, The Rundown reportedVERIFIEDTRUEBENCHMARKtherundown.ai
  16. 16Claude Opus 5 scored 74.0 on DeepSWE v1.1, The Rundown reportedREPORTEDMOSTLY TRUEBENCHMARKdeepswe.datacurve.ai
  17. 17GPT-5.6 Sol scored 73.0 on DeepSWE v1.1, The Rundown reportedREPORTEDMOSTLY TRUEBENCHMARKopenai.com
  18. 18DeepSeek's changelog lists DeepSeek V4.1 Flash scoring 90.6 on Terminal-Bench 2.1VERIFIEDTRUEBENCHMARKapi-docs.deepseek.com
  19. 19DeepSeek's changelog lists DeepSeek V4.1 Flash scoring 30.0 on Terminal-Bench 3.0VERIFIEDTRUEBENCHMARKapi-docs.deepseek.com
  20. 20V4.1 Flash Max Effort's LiveBench category scores run from a low of 70.0 to a high of 93.3.REPORTEDMOSTLY TRUEBENCHMARKCORRECTED IN COPYstraitstimes.com
  21. 21DeepSeek V4.1 Flash is a 552-billion-parameter mixture-of-experts modelVERIFIEDTRUEMODELdeepseek.com
  22. 22DeepSeek V4.1 Flash activates only about 8 billion parameters to read input and 16 billion to write outputVERIFIEDTRUEMODELhuggingface.co
  23. 23Yotta Labs puts the DeepSeek V4.1 Flash checkpoint at about 510 GB, three times the size of V4 FlashREPORTEDMOSTLY TRUEATTRIBUTIONyottalabs.ai
  24. 24DeepSeek V4.1 Flash cache hits cost $0.003 per million input tokens off-peakVERIFIEDTRUEPRICEdeepseek.ai
  25. 25The LiveBench gap between Anthropic's best model (83.4) and DeepSeek V4.1 Flash (81.1) is 2.3 points, about 2.8%COMPUTEDCOMPUTED
  26. 26A job of 1 billion input tokens and 100 million output tokens on DeepSeek V4.1 Flash costs about $210 at off-peak ratesCOMPUTEDCOMPUTED
  27. 27A job of 1 billion input tokens and 100 million output tokens on DeepSeek V4.1 Flash costs $420 at peak ratesCOMPUTEDCOMPUTED
  28. 28A 20-step agent that succeeds 99% of the time per step finishes about 82% of runsCOMPUTEDCOMPUTED
  29. 29A 20-step agent that succeeds 98% of the time per step finishes about 67% of runsCOMPUTEDCOMPUTED
  30. 30DeepSeek V4.1 Flash's DeepSWE v1.1 margins are 0.2 points over Claude Opus 5 and 1.2 points over GPT-5.6 SolCOMPUTEDCOMPUTED
  31. 31The China-U.S. AI gap narrowed 12 points in roughly eight months in 2026COMPUTEDCOMPUTED
  32. 32The fall in GPT-4-class pricing from $30 to $0.40 per million tokens is a 98.7% drop, or 75 times cheaper, in about three yearsCOMPUTEDCOMPUTED
  33. 33Repeating a 75-fold price drop by 2031 would make GPT-4-class capability cost about half a cent per million tokensCOMPUTEDCOMPUTED
  34. 34The projected $255 billion AI inference market in 2030 is about 2.5 times larger than the $103 billion in 2025COMPUTEDCOMPUTED
  35. 35The article's example golden set uses 200 requests pulled from last month's logsCOMPUTEDCOMPUTED

Every claim listed here was extracted from this article and checked against live sources before publication. The verdict is the checker's, not the writer's. Claims the check removed are counted but not republished.

Audit receipt KODA-20261005-220E39C12200
Filed underStrategyDeep Dive05 October 2026
Browse the Deep Dive archive

Get the morning Signal

192 editions so far, one a day. Unsubscribe anytime.