K Koda Intelligence
exploreDeep Dive
DEEP DIVE BRIEFING № 128 · 02 August 2026
Live Intelligence Fact-checked

OpenAI is making tokens nearly worthless
on purpose

On July 31, 2026, OpenAI cut GPT-5.6 Luna's price by 80 percent to $0.20 per million input tokens, and framed it as "abundant intelligence." The company says cost per unit of a given intelligence level has fallen roughly 40x per year. ChatGPT's Work agent and GPT-Live shipped inside the same 48-hour window. That cadence is the real signal: a platform that can absorb your roadmap in a quarter.

7 MIN READ · BY THE KODA EDITORIAL TEAM · STRATEGY · PLATFORM ECONOMICS
$0.20LUNA INPUT↓ OPENAI PRICING
80%PRICE CUT↓ JULY 31 2026
$1.20LUNA OUTPUT↓ PER M TOKENS
keyboard_arrow_down
graphic_eq
LISTEN · AUDIO BRIEFINGThe conversation · ~2 min
smart_display
WATCH · VISUAL NARRATIVEAnimated breakdown · ~2 min
play_arrowPLAY · YOUTUBE
LUNA INPUT$0.20↓ OPENAI PRICING PRICE CUT80%↓ JULY 31 2026 LUNA OUTPUT$1.20↓ PER M TOKENS TERRA INPUT$2.00· 10X PRICE BAND COST DECLINE40X / YR↓ OPENAI ESTIMATE SHIP WINDOW48 HOURS↑ WORK + GPT-LIVE COMMITTED COMPUTE30 GW↑ ALTMAN TCO$1.4T↑ UTILITY SCALE

On July 31, 2026, OpenAI cut the price of GPT‑5.6 Luna by 80 percent. It now costs $0.20 per million input tokens and $1.20 per million output tokens. An 80 percent cut means the old price was five times higher. Same model family. One announcement. Five to one.

OpenAI did not bury the reasoning. The post was titled "Building abundant intelligence," and it defined the goal in plain language: intelligence that "keeps getting more capable, more affordable, and more valuable to the people who use it." In its own recommendations writing, OpenAI states that the cost per unit of a given level of intelligence has fallen roughly 40x per year in recent years. Sam Altman has said the long-run price should converge toward the cost of electricity, and that a typical ChatGPT query burns about 0.34 watt-hours.

I cannot tell you whether that cost curve holds for another three years. Nobody can. But the price cuts are not generosity and they are not a Black Friday sale. They are a deliberate move to make the input you resell nearly worthless, and the products built thinly on top of it optional. Here is how to read the strategy, and where to stand so a price cut helps you instead of ending you.

The Floor, the Furniture, and the Foundation

Everything in an AI-native product sits in one of three tiers.

PRICE LEDGER · JULY 2026OPENAI PRICING · ALTMAN ON X · TECH POLICY PRESS

Four numbers that describe a floor being pushed toward zero.

Luna price cut OpenAI · July 31, 2026 announcement
80%
Annual cost decline OpenAI · cost per unit of intelligence
40x
Sol Fast mode speed OpenAI · 2x price, no intelligence change
2.5x
Committed compute TCO Altman · roughly 30 gigawatts
$1.4T

The Floor is raw intelligence per token. It travels in one direction, and OpenAI just told you which. Luna at $0.20 input, Terra at $2.00 input: a 10x price band between two models in the same generation, and both ends keep dropping.

The Furniture is what most 2024 and 2025 startups actually built. A prompt library, a nice interface, a system message, a retrieval step. Furniture has value only while the room is expensive and empty. The day the platform ships the same feature natively, the furniture goes out with the trash.

The Foundation is the part a price cut cannot reach. Proprietary data nobody else can collect. Ownership of a workflow from trigger to outcome. Distribution into an audience that trusts you. Liability you are willing to carry that a general model vendor will not.

Test yourself with OpenAI's own unbundling. GPT‑5.6 Sol now offers a Fast mode at up to 2.5 times the speed for twice the price, with no change in intelligence. Speed, reliability, and cost are separate dials now, sold separately. If your entire pitch was "we picked the right model for you," that pitch is now a dropdown menu.

The Utility Endgame, and Who Survived the Last One

Zoom out past the news cycle. The pattern is old, and its shape is knowable.

Trying to move towards a true platform, where people and companies building on top of our offerings will capture most of the value.· SAM ALTMAN ON X · 2026

Scarcity buys margin. Abundance buys volume. Every general-purpose input in industrial history has walked from the first to the second, and the walk destroys whoever priced their business on the scarcity. OpenAI's June 8, 2026 plan, written by Altman and Pachocki, opens with rural American electrification in the 1920s. That analogy is a tell. Nobody built a durable business selling access to electricity. They built businesses selling refrigeration, radio, and light at night.

Altman has been explicit about who is supposed to profit. On X he wrote that OpenAI is "trying to move towards a true platform, where people and companies building on top of our offerings will capture most of the value," and that the eventual goal is an AI cloud enabling huge businesses. Read that as an invitation and as a warning. The invitation is real. The warning is that the platform decides where the floor sits, and you do not.

The capital backs the intent. Altman has described roughly 30 gigawatts of committed compute at a total cost of ownership near $1.4 trillion. Utility scale, not software scale. You do not spend that to protect a premium price. You spend it to push marginal cost so low that competitors cannot follow and customers cannot justify leaving.

Costco has sold the same hot dog and soda for $1.50 since 1985. The hot dog is the floor. The membership is the foundation. Anyone who tried to beat Costco on hot dogs misread the business entirely.

Now the counterposition, because the commoditization thesis has serious critics. Writing in Tech Policy Press in March 2025, Trent Kannegieter argued that model commoditization "hasn't arrived yet" even if it is a real possibility. Other analysts at the same publication have called OpenAI's industrial policy writing a "policymercial," a blend of policy argument and corporate marketing, and faulted the superabundance frame for obscuring who actually controls scarce compute and energy.

That critique lands. Abundance inside one vertically integrated stack is not commoditization. It is platformization. Satya Nadella's Jevons paradox post from January 2025 made the incumbent case cheerfully: as AI gets cheaper, usage explodes, and the platform captures the volume. Cheap and fungible are different words. Tokens are getting cheap. Agents, memory, tool permissions, and enterprise identity are not becoming fungible at all.

My read is that both things are true, and the combination is what should change your plan. The unit price of intelligence collapses. The switching cost of the surrounding stack rises. Anyone whose gross margin depends on the first while ignoring the second is running a business with an expiration date they did not set.

The shipping cadence is the real signal. ChatGPT's Work agent and GPT‑Live landed inside the same 48-hour window. A company that can ship two vertically integrated products in two days can absorb your feature roadmap in a quarter. Maybe that pace is a burst tied to one release cycle rather than a steady state. Plan as if it is steady.

Beginner's mind helps here, the shoshin idea of dropping what you think you know. The 2024 playbook was "wrap the model." The 2026 playbook is "own the thing the model cannot see."

Three signals inside the same shift

FURNITURE RISK
80%

A single announcement erased five to one on price.

GPT-5.6 Luna went to $0.20 input and $1.20 output per million tokens in one post. Businesses whose gross margin lived on token markup lost their spread overnight. Same model family, one announcement, five to one.

UNBUNDLING
2.5x

Speed, reliability, and cost are separate dials now.

GPT-5.6 Sol sells a Fast mode at up to 2.5 times the speed for twice the price with no change in intelligence. If your pitch was picking the right model for a customer, that pitch is now a dropdown menu.

PLATFORM PACE
48 HRS

Two vertically integrated products in two days.

ChatGPT's Work agent and GPT-Live landed inside the same 48-hour window. A company shipping at that cadence can absorb a startup's feature roadmap in a quarter. Plan as if the pace is steady, not a burst.

2031

Run the arithmetic forward, gently. If cost per unit of a given intelligence level keeps falling at anything close to the 40x annual rate OpenAI cites, then the token cost inside your product in five years rounds to zero on your P&L. Not cheap. Zero. Price your 2031 business as if inference is free and see what is left standing.

What is left is the asymmetric bet. Data you own compounds. Workflow ownership compounds. Distribution compounds. Model access does not, because it resets to the platform's price every few months.

OpenAI has stated internal targets for an intern-level automated AI research assistant around September 2026 and a genuine automated AI researcher by March 2028. Even if both slip by two years, the direction is set. Continuous, always-on agents replace single-call assistants, and that only works economically at Luna-level pricing.

Here is the contrast pair worth taping to your monitor. Renting intelligence buys you a demo. Owning a workflow buys you a company.

The risk framing matters more than the optimism. If you are Foundation and the price collapses, you win twice: your unit economics improve while your moat stays put. If you are Furniture and the price collapses, you lose twice, because your margin evaporates and your differentiation ships as a native feature. Same event. Opposite outcomes. That asymmetry is the entire decision.

What to Build This Weekend

Stop theorizing and get your reps in. The goal this weekend is one small thing that would survive a 90 percent price cut on tokens.

First, pick a workflow you personally own, not a general capability. Marketing reporting is a good starting point. Connector.wtf is a free, read-only MCP server that exposes Google Ads, Meta Ads, and LinkedIn Ads data straight to ChatGPT or Claude. MCP means Model Context Protocol, a standard way to hand a model access to a live data source. The value there is your account data and your judgment, not the model.

Second, put the agent where the work already happens. Claude Tag sits resident inside a Slack channel with shared context, so follow-up questions do not start from zero. That is workflow ownership in its cheapest possible form.

Third, if the task touches local files or private systems, try Manus My Computer, the new macOS and Windows desktop app that runs the agent on your machine instead of in the cloud. Local execution is a Foundation move. It creates a data boundary that a general platform cannot casually reach across.

Fourth, ship a rough interface. Architect.new surfaced this week on There's An AI For That, in the current wave of prompt-to-app builders. Use it for the ugly version that works, not the beautiful version that does not.

Expect breakage. Your first MCP connection will fail on auth. Your first Slack agent will answer confidently and wrongly. That is normal, and you do not need a CS degree to fix it, only patience and a habit of testing on real data. Take it step by step, build one tiny thing, then ask the only question that matters: if intelligence were free tomorrow, would anyone still pay me? Build the version where the answer is yes.

DOJO · BUILD THIS WEEKEND

Build one small thing that survives a 90 percent price cut on tokens.

  1. Pick a workflow you personally own. Not a general capability. Wire marketing reporting through Connector.wtf, a free read-only MCP server that exposes Google Ads, Meta Ads, and LinkedIn Ads data to ChatGPT or Claude. The value is your account data and your judgment, not the model.
  2. Put the agent where the work already happens. Try Claude Tag resident inside a Slack channel with shared context so follow-up questions do not start from zero. That is workflow ownership in its cheapest possible form.
  3. Create a data boundary, then ship the ugly version. Use Manus My Computer on macOS or Windows to run the agent locally when the task touches private files, then build a rough interface with Architect.new. Expect auth failures and confident wrong answers on the first pass.
Train the full skill in The Dojoarrow_forward
THE BOTTOM LINE

Renting intelligence buys a demo. Owning a workflow buys a company.

The unit price of intelligence is collapsing while the switching cost of the surrounding stack rises, and both facts are true at once. Critics like Trent Kannegieter are right that abundance inside one vertically integrated stack is platformization, not commoditization, which makes the platform's control of the floor the real risk. If you are Foundation and the price collapses, your unit economics improve while your moat holds. If you are Furniture, your margin evaporates and your differentiation ships as a native feature. Price your 2031 business as if inference is free, then ask whether anyone would still pay you.

EDITORIAL RECEIPTKODA-20260802-5682D81A3880
As of02 August 2026MethodClaim extraction, dated-evidence review, and temporal consistency gate.CorrectionsContact the Koda desk
Filed underStrategyDeep Dive02 August 2026
Browse the Deep Dive archivearrow_forward

Want this every morning?

AI analysis, world news, markets, and tools. One briefing, delivered free.

One email per day. No spam. Unsubscribe anytime.