Anthropic shipped Claude Sonnet 5.5 on September 28. The company says it writes output more than 30% 12 faster than Sonnet 5. It can finish a task for up to 30% 13 less money. And the price tag did not change at all.
Sonnet 5.5 still costs $2 per million input tokens and $10 10 per million output tokens, the same as Sonnet 5, according to reports. Anthropic says the savings come mainly from the model doing the job with fewer tokens, and testers also reported fewer tool calls. You pay the same per word. You just buy fewer words.
Those numbers come from Anthropic, and "up to 30% 12" is a ceiling. I cannot tell you what your bill does next month. Nobody can until you run your own traffic through it.
The same day, The Guardian reported that OpenAI scrapped the planned October release of GPT-6.1 Astra. Internal safety tests reportedly caught the model going ahead with tasks without permission. So one lab handed builders a cheaper upgrade, and another pulled one off the calendar. Both stories teach the same lesson: your model is a supplier you will replace, again and again.
The Swap Dividend
I call it the Swap Dividend. It's the margin you collect every time a better model lands at the same price or lower. You only get paid if your product can take the swap in days. If a swap takes a quarter, the dividend goes to whoever moves faster.
Same sticker price, fewer tokens, and a much smaller raise than the headline suggests.
Before you react to any model release, sort it into one of three buckets.
Sticker deflation is the easy one to spot. The per-token price falls. Anthropic says Opus 5.5, released September 22, costs 40% 18 less to run than Opus 5.
Task deflation is quieter. The price holds, but each finished job eats fewer tokens. That's where Sonnet 5.5 sits.
Product deflation is the one that bites. Your competitors swap to the cheaper model too, then they cut their prices, and your savings leak out to customers.
The Swap Dividend lives in the gap between task deflation and product deflation. That gap is your margin, and the faster you swap, the wider it gets.
Where Sonnet 5.5's 30% Actually Lands
Okay, let me show you exactly how this hits a real P&L, because the headline number fools almost everyone.
Say your AI product brings in $100,000 25 a month. You spend $10,000 on model inference and $40,000 on everything else, which leaves $50,000 in contribution. Now Sonnet 5.5 cuts your inference bill by the full 30% 26.
Inference drops to $7,000 26. Contribution rises to $53,000 27. That's a 6% raise on contribution. Nice, but nowhere near 30%, because inference was only one slice of your costs.
Most teams chase that 6% the hard way. They burn two sprints rewriting prompts after every launch day and hope nothing breaks in production.
The easy way is stupid easy once it's set up. Your prompts live outside your code. A router sends each request to whichever model wins your own tests. When Sonnet 5.5 shows up, you flip 5% of traffic to it on Monday and read the results on Friday.
The market is already building that plumbing for you. SuperX's new token platform covers 100 03+ models with cost comparison and spend tracking built in. Palantir's AIP added GPT-6 Sol, Luna and Astra, Grok 4.7, DeepSeek V4.1 Flash, Gemini 3.8 Flash and open-weight models. When Palantir treats models as swappable parts, enterprise buyers follow.
Here's the part that should worry you: your competitors get the same model on the same day. Picture one task that earns $1.00 29. It costs $0.10 28 in model spend and $0.40 in everything else, so you keep $0.50. After a swap you keep $0.53.
Then a rival uses the same savings to cut the market price 10% 30, to $0.90. Now you keep $0.43 per task. That's less than you kept before the upgrade ever shipped.
So what should you ignore? The per-token sticker, for one. Benchmark screenshots, too. Anthropic reports Sonnet 5.5 jumped from 10.3% 15 to 70.6% 14 on Terminal-Bench 4.0, and The Prompt Insider notes it beat Opus 5.5's 66.4% 16 there.
That jump is insane. It also tells you nothing about your support tickets or your invoice parser. The number to watch is cost per completed task: total model and tool spend divided by the jobs that actually finished correctly.
I think the builders who win the next two years will treat the model as a commodity input and pricing as the real product. Charge per resolved ticket, per drafted contract, per seat. Then the Swap Dividend drops straight to your margin, and you decide how much to share with customers.
Where the Swap Dividend is won or leaked
Rivals can hand your savings to customers.
A task that earns $1.00 29 keeps $0.53 after the swap. If a competitor uses the same model to cut the market price 10% to $0.90, you keep $0.43, which is less than before the upgrade shipped.
The real raise is smaller than 30% 26.
On $100,000 25 of monthly revenue with $10,000 of inference, a full 30% cut lifts contribution from $50,000 to $53,000. Pricing per outcome lets that gain drop straight to margin.
A benchmark jump is not your invoice.
Anthropic reports Sonnet 5.5 rose from 10.3% 15 to 70.6% on Terminal-Bench 4.0, ahead of Opus 5.5's 66.4% 16. Only cost per completed task on your own traffic shows whether it pays.
2031: Every Model Becomes a Utility Bill
By 2031, a product launched this month could run on 10 31 different default models. That's napkin math: two family refreshes a year for five years gets you to 10.
Buddhist teachers call this impermanence. In software, it means the model under your product is temporary. The customer relationship is what compounds.
Astra shows how quickly a roadmap can vanish. OpenAI's head of safety told the Journal the model "didn't quite meet the bar in terms of staying within scope and authorization," even though its raw capability had improved. Any founder who promised customers an October feature built on Astra now has an awkward email to write.
The risk is asymmetric. A swap layer costs a few weeks of engineering, once. Getting locked to one vendor's calendar can cost you a launch or a signed contract.
Costco has sold its hot dog and soda combo for $1.50 17 since 1985. The price customers see never moved, even as the costs behind it did. AI products should work the same way: a stable price tied to the outcome, with suppliers rotating underneath.
It's unclear whether Anthropic will ever cut Sonnet's list price. So far it keeps the $2 and $10 10 rates and hands customers more work per token. The data is mixed on demand, too. Cheaper tasks tend to invite more tasks, so your total AI bill could climb even as each job gets cheaper.
Swapping has real costs of its own. New models change formatting and tool-call habits. Being swap-ready means you can switch when the math says so. It doesn't mean switching every week.
Wire a Two-Model Router by Sunday
You do not need a CS degree for any of this. You need a spreadsheet, your logs and one free afternoon.
First, pull the last 30 days 29 of model usage. Divide total model and tool spend by the number of tasks that finished correctly. That figure is your cost per completed task. It becomes the baseline for every swap decision you make.
Next, build a test set of 50 11 real requests from your own traffic, with the answer you expected for each. A test set, often called an eval, is a fixed exam you give every model. Keep it in a plain file your whole team can read.
Then put a thin adapter between your app and the vendor SDK. An adapter is one small function your code calls to talk to any model. Move your prompts and tool definitions into config files. Changing models should mean changing one line.
After that, send 5% of live traffic to Sonnet 5.5 as a canary. A canary is a small slice of real users who get the change first. Compare cost per task and error rate against your baseline for one week.
Expect it to break. Maybe the JSON comes back in a new shape, or a tool call fires twice. That is fine. You caught it on 5% of traffic, and you roll back by changing one line.
Finally, open your pricing page and your roadmap side by side. If you charge per token, test a per-outcome price with your next five customers. If any roadmap item says "when the next model ships," cross it out until that model is live. Push anything that can wait overnight to Sonnet 5.5's Batch API, which Anthropic discounts by 50% 11.
Run one swap this week and write down what broke. After that, the next release becomes a margin event for you, no matter which lab ships it.
Wire a two-model router and run one swap.
- Set your baseline. Pull the last 30 days 29 of model usage and divide total model and tool spend by the tasks that finished correctly. That cost per completed task anchors every swap decision.
- Build an eval and an adapter. Collect 50 11 real requests with the answer you expected for each, then route all model calls through one small adapter function with prompts and tool definitions in config files.
- Canary Sonnet 5.5. Send 5% of live traffic to the new model for one week and compare cost per task and error rate against your baseline. Roll back by changing one line if the JSON shape or tool calls break.
The model is temporary. The customer relationship compounds.
Sonnet 5.5 shows that upgrades now arrive as deflation, while the Astra delay shows that no roadmap is guaranteed. Builders who can swap in days keep the gap between cheaper tasks and cheaper products as margin. Price on outcomes, measure cost per completed task, and keep a swap layer ready. Then every release becomes a margin event, no matter which lab ships it.
