DeepSeek's smallest new model scored 81.1 01 on LiveBench in September. Anthropic's best model scored 83.4 02. The gap is 2.3 25 points, about 2.8%.
The model is also cheap. DeepSeek's rate card lists DeepSeek-V4.1-Flash at $0.15 05 per million cache-miss input tokens ($0.003 per million on cache hits) and $0.60 per million output tokens off-peak. A job of 1 billion 26 input tokens and 100 million output tokens costs about $210. At peak rates it costs $420 27.
We should own the weak spot up front. Nobody can say for sure that a 2.3 25-point gap means anything. It may be noise.
If capability keeps converging at that pace, the model you pick today is a supplier you will re-bid every quarter.
Cost Per Successful Task Is the Unit
A leaderboard tells you how a model did on someone else's exam. A builder pays for something narrower: outputs that pass and ship. Several 2026 analyses of the DeepSeek release settle on the same yardstick for that: cost per successful task. Take everything you spent to get answers and divide by the answers you accepted. That spend counts retries and human review too.
What a 1-billion-token job costs on V4.1 Flash
Every fact in this story sorts into one side of that fraction. The top half is spend. DeepSeek's off-peak input price is $0.15 05 per million tokens, with cache hits at $0.003. Peak hours double every rate, according to DeepSeek and The Rundown. Those hours run 01 07:00 to 04:00 and 06:00 to 10:00 UTC on weekdays.
DeepSeek's launch post says V4.1 Flash's KV cache needs a quarter of the high-bandwidth memory and an eighth of the SSD storage of the prior generation. The KV cache is the model's working memory of the conversation so far, and agents reread it constantly. DeepSeek says cache charges often make up a large share of agent costs.
The bottom half is accepted work. That is where cheap models can lose. Errors compound across agent steps. A 20 28-step agent that succeeds 99% of the time per step finishes about 82% of runs. At 98% 29 per step it finishes about 67%.
So model choice behaves like a price-performance setting. When two models land the same share of accepted outputs on your tasks, the cheaper one wins the fraction. When their pass rates differ, the leaderboard gap stops mattering, and your own numbers decide.
Why a 2.3-Point Gap Is Thin Evidence
The Flash label is the first sign that something has shifted. DeepSeek calls V4.1 Flash the smallest model in its new architecture family. It is a 552 21-billion-parameter mixture-of-experts model that wakes only about 8 billion 22 parameters to read input and 16 billion to write output. A mixture-of-experts model sends each token to a few specialist sub-networks, so most of the model sits idle on any given step.
Still, the tie on paper is shakier than the headline suggests. One 2026 benchmark comparison estimated that with about 1,436 13 test items, gaps under roughly 2.7 points can be treated as noise. The DeepSeek-Anthropic gap is 2.3 25. It is unclear whether the ranking between the two would survive a rerun.
DeepSeek's own numbers show the same pattern. At maximum effort, V4.1 Flash scored 74.2 15 on DeepSWE v1.1, against 74.0 16 for Claude Opus 5 and 73.0 17 for GPT-5.6 Sol, The Rundown reported. Those margins are 0.2 30 and 1.2 points. DeepSeek's changelog also lists 90.6 18 on Terminal-Bench 2.1 and only 30.0 19 on Terminal-Bench 3.0.
Shape matters more than the average. V4.1 Flash Max Effort's LiveBench category scores run from a low of 70.0 20 to a high of 93.3. Your product lives in one or two of those categories. The aggregate only earns the model a trial on your tasks.
The sharpest case study came from DeepSeek itself, in the same week. Its September 9 launch notice said every deepseek-v4-pro request would route to V4.1 Flash from 04 09:00 UTC on September 14. Then its changelog reversed course "in response to user demand," keeping V4 Pro with billing unchanged. Yotta Labs notes there is still no date for a V4.1 Pro.
Swapping, in other words, runs in both directions. Providers change what sits behind a model name, sometimes within a week. A team that pins versions and reruns its tests notices. A team that hard-codes one provider finds out from its customers.
Cash out the door is the honest measure here. A token price is a list price. Retry calls and reviewer hours are what land on the invoice. A cheap model that pushes more cases to humans can cost more per accepted output than a premium one.
Jurisdiction is the other filter. DeepSeek is a Chinese company, and for buyers bound by procurement rules or data-residency laws, that outweighs any token price. The MIT-licensed weights on Hugging Face allow self-hosting. Yotta Labs puts the checkpoint at about 510 GB 23, three times V4 Flash, which changes the hosting math.
The evidence supports one decision today. Build so the model is a setting, and send high-volume, checkable work to whichever model clears your tests cheapest. Three things would change the call: your tests failing on high-cost cases, a U.S. release reopening the gap past the noise band, or DeepSeek raising prices. LiveBench refreshes its questions about every six months, so the next snapshot could look different.
Why the leaderboard tie needs your own tests
The LiveBench gap sits under the noise band.
One 2026 comparison estimated that with about 1,436 13 test items, gaps under roughly 2.7 points can be noise. The DeepSeek-Anthropic gap is 2.3 points, so the ranking may not survive a rerun.
Small per-step misses sink long agent runs.
A 20 28-step agent at 99% per step finishes about 82% of runs. At 98% per step it finishes about 67%, and every failed run adds retries and reviewer hours to the bill.
A model name can change meaning within a week.
DeepSeek planned to route every deepseek-v4-pro request to V4.1 Flash from September 14, then reversed course in response to user demand. Teams that pin versions and rerun tests catch these moves early.
2031. The convergence has a slope, and the slope is what matters strategically. Bloomberg Intelligence measured the gap between Chinese and U.S. models' benchmark scores at about 15% 14 early in 2026 and about 3% by September. That is 12 31 points in roughly eight months. At that pace, the top tier sits inside the 2.7 13-point noise band within months, and a leaderboard can no longer tell the models apart.
Price has a longer record. One 2026 analysis tracks GPT-4-class capability from about $30 32 per million tokens in March 2023 to about $0.40. That is a 98.7% drop, or 75 times cheaper in about three years. Repeat it once more by 2031 and the same capability costs about half a cent per million tokens.
Cheap tokens have not shrunk the bill. The same analysis sizes the inference market at $103 billion 34 in 2025 and projects $255 billion by 2030, about 2.5 times larger. If unit prices fall while total spend rises, volume is growing faster than price is falling. More agents and longer contexts absorb the savings.
That math sets up an asymmetric bet. A model gateway is a one-time cost, measured in days of engineering. Lock-in is a cost that grows with every token you send, and the token count is heading up. If convergence continues, each cheaper model becomes a config change. If it reverses, you have spent a few days.
We think the durable asset by 2031 is your own evaluation set. A 2026 analysis argues model capability is becoming less defensible as a standalone moat as prices fall and gaps narrow. Every failed case you log makes the next switch faster and safer, and that compounds. The model becomes a supplier you re-bid. The tests and customer data stay yours.
Wire a Two-Model Router With Evals
Step one: build the gateway. A gateway here means one function that every model call passes through, reading the model name from a config file. Keep all provider-specific code inside it. Point it at your current model first so nothing changes for users.
Step two: build a golden set. A golden set is a fixed list of real inputs with known good answers. Pull 200 35 requests from last month's logs, including the messy ones that broke things before. Write a pass-or-fail check for each one, such as valid JSON with correct field values.
Step three: run both models on the set. Call your current model and deepseek-flash, with versions pinned. Log input tokens, output tokens, latency and the pass-or-fail result for every case. Some cases will fail on both models, and those failures are the most useful rows in the file.
Step four: price the results. Multiply tokens by each provider's rate, then add reviewer minutes at your hourly cost for every failed case. Divide by passed cases. That number is your cost per successful task, and it drives the routing decision.
Step five: route by task type. Send a task type to deepseek-flash only where its pass rate matches your current model and your data rules allow a Chinese provider. Schedule its batch jobs outside 01 08:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, when DeepSeek charges half. Keep high-cost cases on whichever model wins them.
Step six: rerun the set every week. A scheduled job that replays 200 cases takes minutes. DeepSeek's September plan to reroute deepseek-v4-pro showed a model name can change meaning within a week. The weekly rerun is how you catch the next change before a customer does.
Wire a two-model router and price it by accepted output
- Build one gateway function. Route every model call through a single function that reads the model name from a config file, and point it at your current model first so users see no change.
- Assemble a 200-case golden set. Pull 200 real requests from last month's logs, including the messy ones, and write a pass-or-fail check for each, such as valid JSON with correct field values.
- Price and route by task type. Run both models with pinned versions, add reviewer minutes for each failure, and divide total spend by passed cases. Send a task type to deepseek-flash only where pass rates match and your data rules allow a Chinese provider, then rerun the set weekly.
Own the tests and re-bid the model every quarter.
V4.1 Flash lands within 2.3 25 points of the top U.S. model at $0.15 05 per million input tokens, and the China-U.S. gap fell from about 15% 14 to about 3% 12 in roughly eight months. A gateway costs a few days of engineering, while lock-in grows with every token you send. Your golden set and failure log compound with each switch. We would build so the model is a setting and let cost per successful task decide.
