Anthropic had a 50% price increase written on its pricing page. Input tokens on Claude Sonnet 5 were going from $2 to $3 per million. Output was going from $10 to $15. The date was September 1, 2026.
On August 10, 2026, Anthropic deleted the plan. The docs now say the $2/$10 rate is the standard price and the scheduled increase will not occur. No hedge. No capacity caveat. No "we hope."
That is not a discount. A discount is a lower number. This was a company putting in writing the money it chose not to collect from customers who had already budgeted to pay it.
Sonnet 5 launched June 30, 2026 with that price labeled introductory through August 31. Teams planned around the step-up. Some were front-loading batch jobs. Some were scoping migrations to dodge the bump. All of that work is now dead, and here is why that matters more than the two dollars.
The Certainty Premium Beats The Discount
Buyers do not pay for the lowest price. They pay for the price they can put in a spreadsheet.
Four numbers that explain why a cancelled increase outsells a price cut.
A vendor can take three postures on price, and they are not equally valuable to the person signing off on your infra budget.
A promo is a low rate with an expiry date attached. Fine for trials, useless for architecture. Nobody builds a two-year roadmap on a footnote that dies in eight weeks.
Ambient pricing has no date and no promise, just "subject to change." That is where most AI pricing lives. It feels safe until the email arrives.
Written pricing is a specific dated increase, publicly cancelled in the documentation. That is what Anthropic did on August 10. The certainty premium is the extra adoption you earn by moving from Ambient to Written.
Here is the tell that this is a deliberate posture and not a fluke. Eight days later, on August 18, Anthropic announced it was extending the 50% boost to Claude Code weekly limits through August 31, in what appeared to be the fourth such extension. The wording was "we hope to make this a permanent change," with a capacity explanation attached. That boost launched May 13, 2026 and has been temporary ever since.
Same company, same month, two completely different commitments. Anthropic will write down, with a date, a price it is not charging you. It will not write down, with a date, whether your coding capacity survives next month. Read that gap carefully. It tells you exactly what they think is scarce.
Why A Cancelled Hike Outsells A Price Cut
Price is not a number. Price is a risk position the buyer is taking on your behalf.
Run the napkin math on a single agent loop. Say one run burns 500,000 input tokens and 50,000 output tokens. At $2/$10 that is $1.00 plus $0.50, so $1.50 per run. At the scheduled $3/$15 it becomes $1.50 plus $0.75, so $2.25.
Now scale it. Ten thousand runs a month is $15,000 today and $22,500 after the hike. That is $7,500 a month of new cost for zero new code. Over twelve months, $90,000. Same prompts, same product, same margin target, worse business.
That is the number a CFO was staring at in July. Cancelling the increase does not make the vendor cheaper than it was yesterday. It removes a $90,000 hole from someone's 2027 model. In offer terms, Anthropic did not lower the price. It lowered the perceived risk of buying, and perceived risk is the most expensive line item in any enterprise deal.
Here is the part most people get backwards. The tokens are the golden eggs. The goose is the agent loop that now lives inside your codebase, your CI pipeline, your retry logic, your evals. Anthropic is not selling inference. It is buying a permanent seat in your architecture, and it is paying for that seat with a foregone 50% markup.
Look at the ladder they built. Opus 4.8 sits at $5 input and $25 output per million, which is 2.5x Sonnet 5 on both sides. Sonnet 5 is the price-performance tier, near-flagship coding and agentic work at mid-tier cost, with a 1M-token context window by default. Freezing the middle rung is what keeps the ladder credible.
The shiny distraction here is chasing the cheapest per-token rate on a spreadsheet. DeepSeek announced a switch to peak and off-peak API pricing taking effect at 16:00 UTC on August 16, 2026, with peak hours costing double off-peak. That can be a real saving if your jobs are latency-insensitive. It is not a saving if your team spends three sprints re-architecting to capture it.
The damaging admission: I do not know Anthropic's inference margin on Sonnet 5, and neither does anyone outside the company. Reported annualized revenue has passed $47 billion, and Anthropic announced a $965 billion post-money valuation in its Series H round on May 28, 2026, with an IPO expected this fall. It is unclear whether a permanent $2/$10 is a healthy margin or a deferred bill. The pricing page says the increase will not occur. It does not say the terms can never change again.
What a written promise costs Anthropic to keep
A dated increase publicly cancelled in the docs.
The pricing page now calls $2 input and $10 output the standard rate and states the scheduled increase will not occur. No capacity caveat, no hedge. That moves Sonnet 5 from ambient pricing to written pricing, which is the only tier a two-year roadmap can rest on.
The Claude Code boost still lives on hope.
On August 18 Anthropic extended the 50% weekly limit boost through August 31, apparently the fourth extension since it launched May 13, 2026. The wording was "we hope to make this a permanent change," with capacity attached. Anything in the Hoped column needs a fallback path.
Nobody outside the company knows the inference margin.
Reported annualized revenue has passed $47 billion and Anthropic announced a $965 billion post-money valuation in its Series H on May 28, 2026, with an IPO expected this fall. Whether permanent $2/$10 is a healthy margin or a deferred bill is unclear. The page says the increase will not occur, not that terms can never change.
2031, When Tokens Are Commodity
Zoom out five years. The per-token price of a mid-tier frontier model is heading toward the cost of bandwidth, and everybody in the room knows it.
When the unit price of a thing collapses, the competition moves to the terms around it. Not the number, the contract. Not the rate card, the reliability of the rate card.
This is classic counterpositioning. A competitor racing toward public markets has structural pressure to show revenue per token going up. A competitor that can absorb flat pricing gets to make a promise the other one cannot comfortably match. Price stability is cheap to announce and expensive to sustain, which is exactly what makes it a moat instead of a coupon.
Notice where the market's attention actually sits. Polymarket has roughly $1.67 million in volume on GPT-6 release timing. Traders are pricing the release cadence, not the rate card. My read on this: the cadence bets are the loud story and the pricing commitments are the quiet one, and quiet ones compound.
The contrast pair to hold onto: a discount buys this quarter's usage, a commitment buys the next three years of architecture. One is a promotion. The other is a switching cost you inflict on yourself voluntarily, which is the only kind that sticks.
The honest caveat is that one model at one company is not an industry regime. Extrapolating from Sonnet 5 to "all AI vendors now compete on stability" is a leap the evidence does not fully support yet. Watch whether the next Sonnet launches without an introductory footnote at all. That is the signal that this hardened into strategy.
Reprice Your Agent Runs Before Friday
You do not need a finance team to do this. You need a spreadsheet and two hours.
First, get cost per run. Pick your three highest-volume workflows. Log input tokens, output tokens, and cached reads for each. Multiply by $2, $10, and $0.20 per million respectively. Cached reads are prompt tokens the provider already stored, billed at a fraction of fresh input.
Second, stress test it. Rerun the same math at $3/$15 and at 2x current volume. If a 50% vendor increase breaks your gross margin, your product is priced wrong, not your vendor. Fix your price before you shop for a cheaper model.
Third, classify every vendor promise you depend on. Make three columns: Written, Ambient, Hoped. The Sonnet 5 rate goes in Written. The Claude Code 50% weekly limit boost, extended four times and currently good through August 31, goes in Hoped. Anything in Hoped needs a fallback path.
Fourth, kill the dodge work. If you scoped a migration purely to escape the September 1 increase, stop it today and reclaim the sprint.
Then spend the reclaimed time on something that compounds. Code Wiki, which Google introduced in November 2025, generates a documentation layer over a codebase, which is the job nobody volunteers for and everybody needs before an agent touches production. MIAPI is worth a scouting pass before you hand-write another integration glue script. Cavya.ai auto-builds glossaries and style guides for localization, useful if your token spend is going to translation prep.
If you are looking at support automation, Enjo.ai claims it autonomously handles 20 to 80 percent of customer tickets. Model your business case at 20 percent. Vendors quote the top of the range and buyers who plan for it get burned.
Expect the first pass to be wrong. Your token logs will double-count something. Your run counts will be off. Get the reps in anyway, because a rough cost model beats an accurate guess, and only cash is real. The rest is accounting.
Reprice your agent runs before Friday.
- Get cost per run. Pick your three highest-volume workflows and log input tokens, output tokens and cached reads. Multiply by $2, $10 and $0.20 per million respectively to get a real per-run number.
- Stress test at the cancelled rate. Rerun the same math at $3/$15 and at 2x current volume. If a 50% vendor increase breaks your gross margin, your product is priced wrong, not your vendor.
- Classify every vendor promise. Build three columns labelled Written, Ambient and Hoped. The Sonnet 5 rate goes in Written, the Claude Code weekly limit boost goes in Hoped, and every Hoped row gets a documented fallback.
A discount buys this quarter. A commitment buys three years of architecture.
Anthropic did not cut a price, it deleted a scheduled 50% increase and put that in writing. For a team running 10,000 agent loops a month, that removes $90,000 from the 2027 model without shipping a line of new code. The trade is obvious once you see it: a foregone markup in exchange for a permanent seat inside your CI pipeline, retry logic and evals. One model at one company is not an industry regime yet, so watch whether the next Sonnet launches with no introductory footnote at all. That is the moment stability stops being a coupon and hardens into a moat.