In January, Chinese AI models handled 11% 03 of the tokens on Vercel. By August, they handled 55% 04. That is a fivefold jump in seven months, according to usage data Vercel shared with CNBC.
OpenRouter shows the same curve. In the week that included Sept. 14 06, Chinese models made up 57% to 67% of its token traffic, CNBC reported on Sept. 26. In February, that share sat between 6% and 13% 07.
The cause is plain. These models are cheap, and they write solid code. DeepSeek V4 Flash was reportedly priced around $0.14 01 per million input tokens until mid-August, while OpenAI's GPT-5.5 sat around $5.00 02. That puts the U.S. model at roughly 36 08 times the price for the same input.
I will admit the weak spot up front: token share is a traffic number. CNBC notes that U.S. frontier models still attract more overall spending. A cheap model can also pile up tokens on low-stakes retries.
The pressure now comes from Washington. Two U.S. House committees are investigating how widely U.S. companies are adopting Chinese models. So the model picker in your stack just grew a new field, and this piece shows you how to fill it in.
The Passport Ledger for Every Call
Every model call now carries three prices. The first is the token price, the number on your invoice. The second is the task price: what it costs to get one change merged, counting retries and human review. The third is the passport price, the risk tied to who built the model and where it runs.
The token price is cheap. It is only one of three prices you pay.
Most teams track only the first. That worked when every serious model came from a handful of U.S. labs. It stopped working in 2026. Harpreet Arora, head of agent infrastructure at Vercel, told CNBC that once Chinese models clear a business's quality bar, the price gap becomes a strong deciding factor.
The passport price is the new column, and it depends on hosting. Open weights means the model files are public, so anyone can download and run them. The same weights can run on a China-hosted API, a U.S.-hosted inference provider, or a server you own. Each setup sends your prompts somewhere different, and the legal exposure varies a lot between them.
The whole ledger fits in one sentence: score every call on all three prices, then route it by the worst one. A cheap token with a heavy passport is an expensive call.
Wiring a Router That Knows Borders
OpenRouter's public rankings, with usage data through Sept. 26 06, show what developers actually run. DeepSeek V4.1 Flash reportedly sits first at 20.3 trillion tokens for the week. Z.ai's GLM 5.3 10 Flash is second at 17.8 trillion, and Tencent's Hy4 preview is third at 11.1 trillion. OpenAI's GPT-5.6 Luna is fourth at 8.6 trillion.
Peter Walker, head of OpenRouter Insights, told CNBC that this year's Chinese open models can now handle advanced agent work, particularly coding. Coding agents are freaking token furnaces. Every loop resends repo context, tool output, test logs and the last patch. A router is the switchboard that decides which model answers each of those calls.
Say your coding agent eats 100 million input tokens a month. At the $0.30 15 per million Alibaba Cloud lists for Qwen3-Coder-Next internationally, that costs $30. At $5.00 02 per million, it costs $500. Push to a billion tokens and the gap becomes $300 against $5,000, before a single output token.
The furnace effect also muddies the headline number. A long-running agent burns many tokens per request, so token share can overstate how many teams actually prefer a model. OpenRouter's request-share table for the Sept. 14 week 13 put DeepSeek at 25.4%, which measures something different from token share. It's unclear how much of the Chinese surge reflects broad preference and how much comes from a few heavy agent loops.
Routers like OpenRouter and Vercel's AI Gateway let you choose which models and providers serve a request, with fallbacks when one fails. Most teams set one default model and walk away. That's like letting air traffic control land every plane on the cheapest runway, whatever cargo it carries.
The fix is small. Tag each request with a sensitivity lane before it hits the router. Map each lane to an allowed list of models and hosts. Then log which model and which host actually answered, every single time.
The real trap is in your prompts. Agents get tuned to one model's quirks, like its tool-call format and output parser. If a provider gets restricted, you migrate under deadline pressure, and those parsers break on day one.
Open weights soften that risk only if you can host them yourself. That takes GPUs and people who understand quantization, which means shrinking a model's numbers so it fits on cheaper hardware. For most small teams, a three-lane router with one tested fallback will beat a twelve-model mesh nobody can explain to the security lead.
Cheap tokens, murky share, rising border rules
The input price gap is too big to ignore.
DeepSeek V4 Flash reportedly ran about $0.14 01 per million input tokens against roughly $5.00 02 for GPT-5.5. At a billion tokens a month, Qwen3-Coder-Next's list price works out to $300 against $5,000 before any output.
Token share may overstate real adoption.
Coding agents resend repo context and logs on every loop, so a few heavy agents can inflate token counts. DeepSeek's request share for the Sept. 14 week was 25.4%, a different measure than its token share.
Federal agencies named six Chinese labs.
The NSA, FBI and CISA advisory cited DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.ai over alleged distillation. The proposed China FIREWALL Act would bar these open-weight models from federal devices. Private use is still legal today.
2031: Provenance Becomes a Procurement Checkbox
Washington moved fast this month. On Sept. 8, the NSA, FBI and CISA issued a joint advisory naming DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.ai over alleged large-scale distillation. Distillation means training a model on another model's outputs to copy its skills. On Sept. 18 14, Rep. Josh Gottheimer announced a bipartisan AI safety agenda.
Separately, the proposed China FIREWALL Act would bar Chinese-developed open-weight models from federal devices. It would also stop agencies from buying software that relies on those models. If you sell to the government, your vendor's vendor now matters. To be clear, these are still proposals and investigations. Private-sector use is legal today.
My read is that by 2031, "which models touch our data, and where do they run?" will sit on every enterprise security questionnaire. Huawei is the precedent. After Washington restricted the company in 2019, small U.S. carriers spent years and federal money pulling its gear out of their networks. The cheap equipment looked brilliant on the invoice and painful on the exit.
The risk is lopsided. Cheap tokens save you money a little at a time, every month. A forced migration costs you all at once, on someone else's schedule. Building the exit ramp takes a few days of work today. That's an easy trade.
Treat every model in your stack as temporary. Google held OpenRouter's weekly request lead for 51 weeks 16 before DeepSeek took it in August. The team that stays portable keeps pocketing savings through every price war. The locked-in team pays whatever the next rulebook demands.
Usage and money also live in different places. A Linux Foundation analysis of OpenRouter traffic from May to September 2025 found closed providers captured about 96% 17 of model-layer revenue. I think the winning posture for most builders is boring: use cheap models hard on low-risk work, and keep proof of where every token went.
Tag Your Router Lanes Before Monday
You do not need a CS degree for this. You need a spreadsheet and one afternoon.
First, export the last 30 days of logs from your router or API dashboards. Sort by model, author and token count. You will probably find a Chinese model already in there, picked as a default or fallback months ago.
Second, sort every workload into one of three lanes. Green covers public or open-source code, tests and docs. Yellow covers internal code with no secrets or customer data. Red covers credentials, customer records, regulated data and anything shipped to a government client.
Third, write the routing rule. Green can go to the cheapest model that passes your evals, wherever it runs. Yellow goes to open weights on a U.S.-hosted provider or your own box. Red stays on providers your security team has approved in writing.
Fourth, pick one fallback model built and hosted in a different jurisdiction. Run 20 real tasks from your own repo through both your default and the fallback. Score cost per merged task: the token bill plus review time, divided by changes that actually ship.
Expect the fallback to break a parser or two. That breakage is the whole point of the test, because you want to find it on a quiet Tuesday. Fix the parser, rerun the 20 tasks, and write down what changed.
Fifth, log provenance on every call: model name, author, host region and date. That one log line becomes your answer when anyone asks where their code went.
Build the ledger one lane at a time. The cheap tokens will still be cheap next week. The difference is that you will know exactly where they are allowed to go.
Give your router lanes and a documented exit ramp.
- Export 30 days of router logs. Sort by model, author and token count, then sort every workload into green, yellow or red lanes based on data sensitivity.
- Test one fallback in another jurisdiction. Run 20 real tasks from your repo through your default and the fallback, then score cost per merged task, including review time. Fix any parsers that break and rerun.
- Log provenance on every call. Record model name, author, host region and date in one log line so you can answer where any piece of code went.
Use cheap models hard, and keep proof of where every token went.
Chinese open models have earned their share of real coding work on price and quality. The open question is whether Washington turns investigations into rules, and on whose schedule. A forced migration costs you all at once, while an exit ramp takes a few days to build now. Tag your lanes, test one fallback and log provenance on every call so the savings survive the next rulebook.
