K Koda Intelligence
exploreDeep Dive
DEEP DIVE BRIEFING № 127 · 01 August 2026
Live Intelligence Fact-checked

The floor under your pricing just melted in 21 days

OpenAI cut GPT-5.6 Luna prices 80% on July 30, 2026, roughly three weeks after general availability. Input fell from $1.00 to $0.20 per million tokens and output from $6.00 to $1.20. Higher-tier Terra took only a 20% trim and flagship Sol did not move at all. Same weights, one-fifth the bill, 21 days later. If your price was cost-plus, that is a threat. If it was value-based, you just got a margin gift.

7 MIN READ · BY THE KODA EDITORIAL TEAM · MARKETS · AI PRICING
80%LUNA CUT↓ OPENAI
$0.20LUNA INPUT↓ FROM $1.00
$1.20LUNA OUTPUT↓ FROM $6.00
keyboard_arrow_down
LUNA CUT80%↓ OPENAI LUNA INPUT$0.20↓ FROM $1.00 LUNA OUTPUT$1.20↓ FROM $6.00 TERRA CUT20%↓ SAME NOTICE SOL INPUT$5.00· UNCHANGED CACHED INPUT$0.02↓ 10X BELOW STANDARD TIME TO RESET21 DAYS· POST-LAUNCH CUT DATEJUL 30· 2026

OpenAI cut the price of GPT-5.6 Luna by 80% on July 30, 2026. The model had been generally available for roughly three weeks. Input dropped from $1.00 to $0.20 per million tokens. Output dropped from $6.00 to $1.20.

Combined, that is $7.00 per million tokens down to $1.40. Same model. Same weights. One-fifth the bill, 21 days later.

GPT-5.6 Terra got a 20% cut in the same announcement, from $2.50 to $2.00 input and, reportedly, $15.00 to $12.00 output. The flagship, Sol, did not move at all. It reportedly still costs $5.00 input and $30.00 output. OpenAI framed the whole thing as "advancing the price-performance frontier," and said Luna now matches models that were frontier-class a year ago for about 6 cents on the dollar per task at nearly 9x the speed.

CNBC reported the cut as a response to cost-sensitive customers and competition from Chinese startups and other tech giants. Cryptopolitan put it more bluntly in its headline: cheap rivals are gaining ground.

Here is the part that matters for anyone with a paying customer. Your cost of goods sold just fell 80% without you doing anything. That is good news. It is also a warning about how you built your price.

The Melting Floor

The floor under your product is melting faster than you can re-sign a customer.

PRICING LEDGER · JULY 2026OPENAI · CNBC · REUTERS · CRYPTOPOLITAN

Four numbers that show where the commoditization actually landed.

GPT-5.6 Luna price cut OpenAI · three weeks after GA
80%
GPT-5.6 Terra price cut Same announcement · $2.50 to $2.00 input
20%
Sol flagship movement Still $5.00 input, $30.00 output
0%
Token line as share of revenue Napkin math · $29/mo plan, day one
7.6%

That is the whole rule. Most SaaS contracts run 12 months. Most AI product pricing gets revisited quarterly, if you are disciplined about it. The input cost underneath it just repriced in three weeks.

Sort your pricing into two buckets. Bucket one is cost-plus: you looked at your token bill, added a margin, shipped a number. Bucket two is value-based: you looked at what the customer saves or earns, and took a slice. Bucket one just got detonated. Bucket two did not notice.

If your price was cost-plus, the melting floor is a threat, because your competitor can now undercut you by 60% and still make money. If your price was value-based, the melting floor is a gift, because your gross margin quietly expanded while your invoice stayed the same. Same event. Opposite outcome. The difference is a decision you made before the cut, not after.

You Are Not Selling Tokens. Stop Pricing Like You Are.

Let me do the napkin math on a real product shape.

If your token cost fell 80% and your business felt materially different, what were you actually selling? If the answer is access, you were reselling somebody else's inventory at a markup, and the supplier just repriced your inventory without asking.· KODA EDITORIAL · JULY 2026

Say you run a document-processing tool. Each customer burns about 1 million input tokens and 200,000 output tokens per month. Before July 30, that cost you $1.00 plus $1.20, so $2.20. After the cut it costs $0.20 plus $0.24, so $0.44. On a $29 per month plan, your gross profit per customer went from $26.80 to $28.56.

Notice how boring that is. Your margin was already 92%. Now it is 98%. The token line was never your business. It was 7.6% of your revenue on day one.

This is the golden goose problem. Founders stare at the golden egg, the API invoice, because it is the number that arrives in an email every month. The goose is the thing that actually pays you: the customer's willingness to pay for an outcome they cannot get elsewhere. Nobody has ever cut the price of your customer relationship by 80% overnight.

So here is the childlike question worth asking out loud. If your token cost fell 80% and your business felt materially different, what were you actually selling? If the answer is "access," you were reselling somebody else's inventory at a markup, and the supplier just repriced your inventory without asking. That is not a business model. That is lipstick on a pig.

Now the damaging admission. I do not know if this is a trend yet. One 80% cut is a data point, not a curve. OpenAI has not made a habit of three-week 80% resets on prior generations, and the flagship tier held firm at $5.00 and $30.00, which suggests premium capability still prices like a differentiated asset rather than a commodity. Reuters framed the move as cuts to the smaller models specifically, not the whole lineup.

The data is mixed on whether this becomes quarterly behavior or stays a one-off shot at cheap competitors. It is unclear whether OpenAI can sustain $0.20 input pricing if inference demand spikes the way elastic demand usually does when a price drops 5x.

But the direction of the arbitrage is not ambiguous. Cost per unit of intelligence is falling. Your LTV:CAC ratio does not care about cost per token at all. It cares about what a customer pays you over their life divided by what you spent to get them. Token deflation does not raise LTV by a cent. It only widens the gap between price and cost, which is exactly the gap a competitor can walk into.

There is a second lever most people ignored in the announcement. Cached input on Luna runs $0.02 per million tokens, another 10x below the new standard $0.20 rate. Architect your prompts so the stable parts get cached, and your effective cost drops again. That is a product engineering decision with a pricing consequence, which is where most real margin lives.

Three signals inside the same shift

MELTING FLOOR
21 DAYS

Cost inputs now reprice faster than contracts renew.

Most SaaS agreements run 12 months and most AI pricing gets revisited quarterly. Luna's input cost repriced from $1.00 to $0.20 in roughly three weeks. Cost-plus pricing built on that floor can be undercut by 60% and the competitor still makes money.

TIERED DEFENSE
$30.00

Premium capability still prices like a differentiated asset.

Luna took the full 80% hit while Terra fell about 20% and Sol held firm at $5.00 input and $30.00 output. Reuters framed the move as cuts to the smaller models specifically, not the whole lineup. Commoditization is hitting the cheap tier first, not the frontier.

ASYMMETRIC BET
2031

Token-hungry architectures get better economics for free.

Multi-step agents, persistent memory and vote-across-models were financially stupid at $7.00 per million combined. At $1.40 they are merely expensive. Cached input at $0.02 per million pushes the effective floor another 10x lower for anyone who architects prompts for reuse.

2031

Zoom out five years. The pattern here is not new, it is just fast.

Cloud compute did this. Storage did this. Bandwidth did this. In every case, the layer that commoditized was the layer closest to the metal, and the value migrated up the stack toward workflow, data, and problem ownership. AWS did not stop being a good business when S3 got cheap. The companies that only resold S3 did.

The asymmetric bet here is obvious once you name it. If token prices keep falling, products designed to consume more tokens per task get better unit economics every quarter, for free. Multi-step agents, persistent memory, aggressive self-checking, running three models and voting: all of those were financially stupid at $7.00 per million combined. At $1.40 they are merely expensive. In 2031 they will be rounding errors.

The flywheel runs like this. Cheaper tokens allow heavier workflows. Heavier workflows produce better outputs. Better outputs justify higher prices to customers, who never see a token count. That is compounding in your favor, and it only works if you decoupled your price from your cost early.

I think the losers of this cut are not the labs and not the enterprises. My read is that it lands hardest on the thin middleware layer, the businesses whose entire margin was the spread between what OpenAI charged them and what they charged you. FourWeekMBA's analysis of the cut made the same point: for real AI-native products, this is a unit-economics reset that makes marginal products profitable. For arbitrage plays, it is an eviction notice.

Salary buys furniture. Equity buys your future. Same logic applies to margin: a cost cut buys you one good quarter, but pricing power buys you the decade.

What to Build This Weekend

Do three small things. Take a deep breath and go in order.

First, write down your gross margin per customer per month, using the July 30 prices. Not your blended margin. One customer, one month, actual token counts from your usage dashboard. If you cannot produce that number in 20 minutes, that is the real finding, and instrumenting it is your weekend project.

Second, run a substitution test. Take your highest-volume workload and route it through Luna at $0.20 input instead of whatever tier you defaulted to. Then evaluate quality on 50 real examples, not vibes. Static model assignments made in July are already stale, and re-running that evaluation should be a recurring calendar item, not a one-time chore.

Third, build one thing that gets cheaper as tokens get cheaper. OfficeCLI 1.0.132 is an open-source single binary that lets an agent create, read and edit .docx, .xlsx and .pptx files directly, which means document workflows that used to need a human middleman can now run end to end. Point an agent at it and see what breaks. Things will break. That is the point of a weekend.

If you want the approvals loop out of your terminal and onto your phone, the new OpenClaw iOS and Android app handles remote control, notifications and task approvals. For a full-stack prototype with a real database behind it, Blink.new is built for first-time builders and will get you to something clickable fast. Lumi.new is newer and thinner on documentation, so treat it as a throwaway project rather than a foundation.

One last note on pricing your own thing. Do not reprice down just because your costs fell. Reprice up if you can show the customer a bigger outcome. The floor is melting. Build higher.

DOJO · BUILD THIS WEEKEND

Instrument your margin before your competitor instruments theirs.

  1. Compute one customer's true gross margin. Use the July 30 prices and actual token counts from your usage dashboard for a single customer over a single month, not a blended average. If you cannot produce that number in 20 minutes, instrumenting it is the weekend project.
  2. Run a substitution test on your heaviest workload. Route it through Luna at $0.20 input instead of whichever tier you defaulted to, then score quality on 50 real examples rather than vibes. Put the re-evaluation on a recurring calendar item, because static model assignments made in July are already stale.
  3. Ship one thing that gets cheaper as tokens get cheaper. Point an agent at OfficeCLI 1.0.132, the open-source single binary that creates, reads and edits .docx, .xlsx and .pptx files, and let it run a document workflow end to end. Things will break, and that is the point of a weekend.
Train the full skill in The Dojoarrow_forward
THE BOTTOM LINE

A cost cut buys one good quarter. Pricing power buys the decade.

One 80% cut is a data point, not a curve, and it is genuinely unclear whether OpenAI can hold $0.20 input pricing once elastic demand arrives. But the direction of the arbitrage is not ambiguous: cost per unit of intelligence is falling, and the value is migrating up the stack toward workflow, data and problem ownership exactly as it did with cloud, storage and bandwidth. The businesses that get evicted are the thin middleware layer whose entire margin was the spread between what OpenAI charged them and what they charged you. For anyone selling an outcome, the token line was never the business anyway, it was 7.6% of revenue on day one. Do not reprice down because your costs fell. Reprice up if you can show a bigger outcome, and build higher while the floor keeps melting.

EDITORIAL RECEIPTKODA-20260801-589B99380589
As of01 August 2026MethodClaim extraction, dated-evidence review, and temporal consistency gate.CorrectionsContact the Koda desk
Filed underMarketsDeep Dive01 August 2026
Browse the Deep Dive archivearrow_forward

Want this every morning?

AI analysis, world news, markets, and tools. One briefing, delivered free.

One email per day. No spam. Unsubscribe anytime.