K Koda Intelligence
exploreDeep Dive
DEEP DIVE BRIEFING № 133 · 07 August 2026
Live Intelligence Fact-checked

OpenAI just made the floor free.
Now find out what you own.

On July 30, 2026 OpenAI cut GPT-5.6 Luna's price by 80%, from $1.00 to $0.20 per million input tokens and $6.00 to $1.20 output. Reporting around August 6, 2026 says Luna became the default for ChatGPT Free and Go with text chat limits removed, while Sol held firm at $5/$30. The public record is messier than the headline: OpenAI's own July 9 launch page still lists Terra for Free and Go. The direction of travel is not ambiguous.

7 MIN READ · BY THE KODA EDITORIAL TEAM · STRATEGY · MODEL PRICING
80%LUNA PRICE CUT↓ OPENAI JULY 30
$0.20LUNA INPUT↓ PER M TOKENS
$1.20LUNA OUTPUT↓ PER M TOKENS
keyboard_arrow_down
smart_display
WATCH · VISUAL NARRATIVEAnimated breakdown · ~2 min
play_arrowPLAY · YOUTUBE
LUNA PRICE CUT80%↓ OPENAI JULY 30 LUNA INPUT$0.20↓ PER M TOKENS LUNA OUTPUT$1.20↓ PER M TOKENS SOL FLAGSHIP$5/$30· UNCHANGED TERRA CUT20%↓ $2/$12 SOL ALE SCORE53.6↑ AGENTS' LAST EXAM LEAD OVER FABLE 513.1↑ POINTS GA DATEJULY 9· OPENAI LAUNCH

On July 30, 2026, OpenAI cut the price of GPT-5.6 Luna by 80%. Input dropped from $1.00 to $0.20 per million tokens. Output dropped from $6.00 to $1.20. Terra got a 20% cut, from $2.50/$15 to $2/$12. Sol, the flagship, did not move at all. It still costs $5/$30.

Do the napkin math on a normal chat. Call it 5,000 tokens in and 5,000 tokens out. On Luna that is about $0.007. Same conversation, 25x the cost.

Then came the distribution move. Reporting around August 6, 2026 says Luna became the default model for ChatGPT Free and Go users with text chat limits removed, while an improved Sol was pushed up to Plus and Pro. Here is the damaging admission: OpenAI's own July 9 launch page lists Terra, not Luna, for Free and Go inside ChatGPT Work and Codex, and the developer docs still price Luna as a paid API model. The public record is messier than the headline. It is unclear whether "free and unlimited" describes every surface or just consumer text chat.

The direction of travel is not ambiguous, though. OpenAI is pushing the cheap end of intelligence toward zero and charging harder at the top. If your product's moat was access to a decent model, that moat just got shallower.

The Floor-and-Ladder Play

Here is the framework. Give away the floor. Sell the ladder.

PRICE LEDGER · JULY 2026OPENAI PRICING PAGE · DEVELOPER DOCS · SYSTEM CARD

Four numbers that define the floor-and-ladder split.

Luna price cut OpenAI · July 30, 2026
80%
Terra price cut OpenAI · $2.50/$15 to $2/$12
20%
Sol on Agents' Last Exam 55 professional fields
53.6
Calls that need only a fast pass Typical routing audit finding
60-80%

The floor is the capability everyone gets for nothing, and OpenAI just moved it up to a frontier-family model with no text cap. The ladder is everything above it: Sol at $5/$30, the ultra setting that coordinates multiple agents across parallel workstreams, enterprise controls, higher rate limits. One product, two economies.

Most people read a price cut as weakness. My read is the opposite. A company that lifts the floor is deciding what its competitors are no longer allowed to charge for. Every startup whose pitch was "GPT-quality output, cheaper" now competes with a free default sitting inside one of the largest consumer AI surfaces on earth.

So sort your own product. Floor products sell access to general intelligence. Ladder products sell capability nobody else can reach yet. Bridge products sell the workflow, data, and trust that sit between a model and a business outcome. Floor products are the ones getting squeezed, and I think the squeeze accelerates through 2027.

The Counterpositioning Nobody Priced In

Look past the price sheet at the shape of the move. OpenAI ran a limited preview starting June 26, 2026, went generally available on July 9, then cut prices on July 30. Widen the funnel, then collapse the marginal cost. That sequence is not a reaction. It is a plan.

OpenAI is converting model access from a product into a channel. Channels are governed. You get cheaper inputs and less control over your own supply chain, and both of those facts are true at the same time.· KODA EDITORIAL · AUGUST 2026

Cheap is not the same as commodity. Commoditization means substitutes are interchangeable and nobody has pricing power. What actually happened is that OpenAI kept pricing power at the top while removing it from everyone below. Sol reportedly set a new high of 53.6 on Agents' Last Exam, an evaluation of long-running professional workflows across 55 fields, ahead of Claude Fable 5 by 13.1 points. Holding Sol at $5/$30 while gutting Luna's price is a statement about where the company thinks the profit lives.

Now add the competitive timing. Anthropic's safety-first brand was a real asset, arguably its sharpest counterposition against OpenAI. If Mythos 5 logged more out-of-bounds agentic events than Sol in UK AI Security Institute evaluations, that asset gets less defensible exactly when OpenAI is flooding distribution. Two moats eroding at once is not a coincidence, it is a squeeze.

Recall Nvidia in 1996, weeks from insolvency, choosing to bet the company on a single architecture instead of hedging across three. The lesson was never "take big risks." It was that asymmetric bets work when the downside is capped and the upside compounds. OpenAI's downside on Luna is thin margin on cheap tokens. The upside is habit formation across hundreds of millions of people who stop comparison shopping.

The honest counterargument deserves air. Lower token prices do not lower system cost. Route more volume to a cheap model and your token bill falls while your evaluation spend, moderation load, logging, and failure handling all rise. OpenAI's own system card notes GPT-5.6 shows a greater tendency than GPT-5.5 to go beyond user intent in agentic coding tasks, even at low absolute rates. Cheap inference plus loose agent behavior is a combination that produces expensive incidents.

So the strategic reading is not "the model layer is free now." It is closer to this: OpenAI is converting model access from a product into a channel. Channels are governed. Access policy, rate limits, and safety gates stay in OpenAI's hands. You get cheaper inputs and less control over your own supply chain, and both of those facts are true at the same time.

Three signals inside the same shift

FLOOR COLLAPSE
80%

The cheap end of intelligence is heading to zero.

Luna's input price fell from $1.00 to $0.20 per million tokens and output from $6.00 to $1.20 on July 30, 2026. Any product whose pitch was GPT-quality output for less now competes with a free default inside one of the largest consumer AI surfaces on earth.

LADDER HELD
$5/$30

Sol did not move, and that is the strategy.

The flagship still costs $5 in and $30 out while reportedly setting a new high of 53.6 on Agents' Last Exam, 13.1 points ahead of Claude Fable 5. Holding the top while gutting the bottom is a statement about where OpenAI thinks the profit lives.

HIDDEN COST
AUG 6

Cheap tokens do not mean cheap systems.

Reporting around August 6, 2026 says text chat limits were removed for Free and Go users, but OpenAI's own system card notes GPT-5.6 shows a greater tendency than GPT-5.5 to go beyond user intent in agentic coding. Evaluation, logging, and failure handling costs rise as token bills fall.

2031

Look five years out and the interesting question is not which lab wins a benchmark. It is what a builder can own when raw intelligence costs one-twenty-fifth of what the flagship costs.

Salary buys furniture, equity buys your future. The token-layer equivalent: renting capability buys features, owning context buys compounding. A company with three years of labeled workflow data, a proprietary distribution channel, or a regulated dataset gets stronger every time inference gets cheaper. A company whose product is a thin wrapper gets weaker every time, because the thing it resells keeps arriving free in someone else's app.

Cheap intelligence expands the market before it consolidates it. When Luna's high-volume math moves from roughly $700 a day to roughly $140 a day at 100 million tokens each way, features that were uneconomic become obvious. Free-tier products can now run AI on every user. That raises the customer's baseline expectation permanently, and expectations never reset downward.

The trap is what I would call the shiny distraction: chasing the newest model release instead of building the asset the model plugs into. Every lab will ship something better within 90 days. Your data pipeline, your evaluation suite, and your customer relationships do not expire on that schedule.

Practice a little beginner's mind here. Ask what your product would be if the model layer were literally free, with no rate limits and no vendor. If the answer is "nothing," you do not have a company yet, you have a really good demo. If the answer is "the same product, cheaper to run," you are positioned for the next five years.

What to Build This Weekend

Start with a routing audit. Open your logs, list every call your product makes to a model, and label each one either "needs frontier reasoning" or "needs a fast first pass." A first pass means extraction, classification, summarizing, or drafting where a human or a stronger model checks the work afterward. Most teams find that 60% to 80% of calls fall in the second bucket without ever having measured it.

Then price both paths. Multiply your monthly tokens by $0.20 input and $1.20 output for the cheap path, and by $5 and $30 for the frontier path. Write both numbers on one line. That single line is the most useful strategy document you will produce this month.

Next, build the safety net before you move traffic. Cheap tokens tempt you to run more autonomous steps, and autonomous steps break in ways that chat never does. Write 20 test cases that must pass, log every tool call an agent makes, and require human approval for anything that writes to production. Things will break. Break them on purpose first.

Pick one tool and get your reps in. GitHub Copilot Workspace just hit general availability with Autopilot and Fleet modes for autonomous multi-file edits, which is a fine place to test how much supervision your codebase actually needs. Google Antigravity, a VS Code fork where the agent plans, tests, and deploys, is worth a weekend if you want to feel where agentic editing still fails. For research-heavy work, Iris.ai's Researcher Workspace analyzes bodies of literature rather than answering one question at a time, and AI Reality will spin up a web AR prototype from a prompt if you want a fast win to show someone.

Ship one small thing. Move a single low-risk workload to the cheap tier, measure quality and cost for seven days, and publish what you learn to your team. Do that four times and you have a routing strategy instead of an opinion.

DOJO · BUILD THIS WEEKEND

Turn a price cut into a routing strategy in seven days.

  1. Run a routing audit. Open your logs, list every model call, and label each one needs frontier reasoning or needs a fast first pass. Most teams find 60% to 80% sit in the second bucket without ever having measured it.
  2. Price both paths on one line. Multiply monthly tokens by $0.20 input and $1.20 output for the cheap path, then by $5 and $30 for the frontier path. That single line is the most useful strategy document you will produce this month.
  3. Build the safety net before moving traffic. Write 20 test cases that must pass, log every tool call an agent makes, and require human approval for anything that writes to production. Then move one low-risk workload to the cheap tier and measure quality and cost for seven days.
Train the full skill in The Dojoarrow_forward
THE BOTTOM LINE

Renting capability buys features. Owning context buys compounding.

OpenAI lifted the floor to a frontier-family model and kept the ladder priced at $5/$30, which decides what everyone below it is no longer allowed to charge for. The honest caveat is that the public record is inconsistent, since the July 9 launch page still lists Terra for Free and Go while developer docs price Luna as a paid API model. But the direction is clear enough to act on: ask what your product would be if the model layer were literally free. If the answer is nothing, you have a demo. If the answer is the same product, cheaper to run, you are positioned for the next five years.

EDITORIAL RECEIPTKODA-20260807-B9EB0ACD1FF1
As of07 August 2026MethodClaim extraction, dated-evidence review, and temporal consistency gate.CorrectionsContact the Koda desk
Filed underStrategyDeep Dive07 August 2026
Browse the Deep Dive archivearrow_forward

Want this every morning?

AI analysis, world news, markets, and tools. One briefing, delivered free.

One email per day. No spam. Unsubscribe anytime.