K Koda Intelligence
DEEP DIVE DEEP DIVE № 204 · 23 September 2026DOCKODA-20260923-BA67917B2777sha-256 of date + article + 24 claims
FILED 23 SEPTEMBER 202624 CLAIMS CHECKED · 10 VERIFIED

Who keeps the eight cents?
Price cuts are now the release headline

OpenAI dropped GPT-6 Sol to $2 input and $10 output per million tokens on September 2204, half the prior $4 and $2002. Anthropic shipped Opus 5.501 the same day at $4 and $20, a 20% token-price cut against Opus 5, while Xiaomi claims up to 20x faster output for MiMo. On one illustrative 20,000-in, 4,000-out task the model bill falls from $0.16 to $0.08. The only question left is who keeps the difference.

7 MIN READ · BY THE KODA EDITORIAL TEAM · MONETIZATION · MODEL PRICING
SOL INPUT$2NOT MEASURED↓ 50%
SOL OUTPUT$10NOT MEASURED↓ 50%
LUNA OUTPUT$0.50NOT MEASUREDFROM $1.20
SOL INPUT$2↓ 50% SOL OUTPUT$10↓ 50% LUNA OUTPUT$0.50FROM $1.20 OPUS 5.5 INPUT$4↓ 20% MIMO SPEED20×XIAOMI CLAIM TASK MODEL COST$0.08↓ 50% CONTRIBUTION$0.92↑ 8pp AMD VALUE$1TSEP 21

A 50% model price cut can raise your margin without adding a customer.

On September 2204, its price fell to $2 per million input tokens and $10 per million output tokens. The preceding Sol rates were $4 and $2002. Tokens are the small text units these services bill for.

That is good news if customers pay for work your product completes. It is dangerous news if they mostly pay for access to a model.

Anthropic and Xiaomi add pressure from different directions. Their claims need care: lower token prices, lower running costs, and faster output measure different things.

The useful question is who keeps the savings. We can answer it with a pricing framework and a task-level margin test.

The Savings Ownership Matrix

I think the strongest business opportunity here is to keep customer value steady while making delivery cheaper. Automatically passing every saving through can waste that opportunity.

PRICE CUT LEDGER · SEPTEMBER 2026VENTUREBEAT · THE NEXT WEB · FORKASTBASE: 24 CHECKED CLAIMS, 4 SHOWN

Four figures that decide whether the cut reaches your margin.

GPT-6 Sol output rate OpenAI · per million tokens, down from $20 NOT MEASURED
$10
GPT-6 Luna output rate OpenAI · per million tokens, down from $1.20 NOT MEASURED
$0.50
Model cost per illustrative task 20,000 input and 4,000 output tokens, down from $0.16 NOT MEASURED
$0.08
Monthly model spend freed Illustrative one million tasks · from $160,000 to $80,000 NOT MEASURED
$80,000

The Savings Ownership Matrix sorts products using two questions: How much of the customer’s workflow do you own? How easily can you change models?

Workflow ownership means your product handles a meaningful piece of the job, such as a completed approval process. A blank chat box gives you much less to defend.

The matrix has four positions:

  • Owners: Strong workflow value and easy model switching. Test cheaper delivery while protecting the result customers already buy.
  • Dependents: Strong workflow value and difficult switching. Invest in testing and portability before assuming the announced savings are yours.
  • Resellers: Weak workflow value and easy switching. Expect competitors to discover the same price cuts. Give customers a reason to stay.
  • Trapped wrappers: Weak workflow value and difficult switching. Reconsider the product before spending heavily on another model migration.

An AI wrapper is typically an application that adds an interface or prompts around another company’s model. It can also add workflow logic or integrations. Being one is a starting point, but it is a fragile destination.

A cheaper GPT-6 Luna does not automatically move your business into the Owners box. You need evidence that customers value something beyond the underlying response.

Let your position guide your next investment. Owners can improve delivery. Resellers need to strengthen the offer before treating lower costs as durable profit.

The Eight-Cent Margin Test

OpenAI provides the clearest starting point. VentureBeat reports that an OpenAI spokesperson confirmed the new Sol and Luna rates are permanent.

A 50% model price cut can raise your margin without adding a customer. That is good news if customers pay for work your product completes. It is dangerous news if they mostly pay for access to a model.· THE KODA ANALYSIS · SEPTEMBER 2026

[The Next Web](https://thenextweb.com/news/openai-gpt-6-sol-luna-api-price-cut) places both models in Codex and ChatGPT Work from September 22, 2026. The cuts reach tools people already use for work.

Sol’s input and output prices both fell 50%. Luna’s input price fell from $0.20 to $0.10 per million tokens.

Luna’s output price fell from $1.20 to $0.50, a reduction larger than half. A blanket “50% cheaper” headline hides differences in what your product actually consumes.

Consider the research brief’s illustrative task: 20,000 input tokens and 4,000 output tokens. Using OpenAI’s reported standard rates gives:

  • GPT-5.6 Sol: $0.08 input plus $0.08 output, $0.16 per task
  • GPT-6 Sol: $0.04 input plus $0.04 output, $0.08 per task

Assume the customer pays $1 per completed task. This is an illustration, not a reported customer contract.

Your contribution before other costs rises from $0.84 to $0.92. The gain is eight percentage points of revenue. The calculation requires neither a new customer nor a higher price.

At the brief’s illustrative volume of one million monthly tasks, model spending falls from $160,000 to $80,000. That leaves $80,000 to allocate elsewhere.

Those savings depend on unchanged task volume and quality holding up. You still need to pay migration costs.

You do not need a complicated dashboard to understand the opportunity. You need a clear invoice and a believable definition of “completed.”

Anthropic requires more caution. [Forkast](https://forkast.news/openais-gpt-6-sol-and-luna-cut-prices-50-and-the-three-tier-family-just-ended-the-coordinated-slowdown/) reports that Opus 5.501 arrived the same day at $4 input and $2002 output per million tokens.

Compared with the research brief’s $5/$25 rates for Opus 5, that is a 20%02 token-price cut. The brief separately attributes a 40% running-cost reduction claim to Anthropic.

It also cites Anthropic’s claim that Opus 5.501 matches Fable 5.1 on most work. It is unclear whether those comparisons use equivalent workloads.

Do not treat those claims as interchangeable discounts. Token pricing is directly visible on the rate card. Running cost depends on how much work the model needs to finish.

Xiaomi introduces another variable. The research brief attributes to Xiaomi a claim of up to 2002× faster output speed for MiMo-V2.6-Pro-UltraSpeed at the same quality.

The supplied material does not establish the test conditions behind that claim. Treat it as a reason to test, not a promised reduction in your complete bill.

The brief also lists MiMo-V2.6-Pro at $0.435 input and $0.87 output per million tokens. At those rates, the same illustrative task costs $0.01218.

That figure covers model charges alone. It excludes the rest of the workflow.

Xiaomi’s MiMo-V2.6 open weights are also available on Hugging Face under an MIT license, according to [Forkast](https://forkast.news/xiaomis-mimo-v2-6-ships-open-weights-at-frontier-class-performance-and-the-timing-is-not-an-accident/). Open weights can let developers run the model themselves, subject to hardware and software compatibility. License restrictions also apply.

Faster generation could improve the economics of self-hosting. It does not establish a new self-hosting cost floor without hardware and operating-cost evidence.

I would measure it this way:

Cost per accepted task = all spending on the workflow ÷ tasks accepted as complete.

Include what you spend on failed attempts and human review. A result that needs repair does not have the same economics as one the customer can use immediately.

The old approach was to pick a model and price around its bill. The easier commercial test starts with one customer result and compares ways to deliver it.

Resist the temptation to name one model the cheapest for every job. Your workload decides.

Three cuts, measured in three incompatible units

MARGIN WINDFALL
8pp

The cheapest growth you will book this quarter.

At $1 per completed task, contribution before other costs rises from $0.84 to $0.92 when the model bill halves from $0.16 to $0.08. That is eight percentage points of revenue with no new customer and no price increase. It holds only if task volume and output quality stay put.

WRAPPER RISK
$0.01218

Weak workflow ownership gets undercut fast.

At the brief's listed MiMo-V2.6-Pro rates of $0.435 input and $0.87 output, the same illustrative task costs $0.01218 in model charges alone. Products that sell access rather than a completed result will meet competitors who found the same rate card. Resellers need a reason to stay before treating lower costs as durable profit.

CLAIM MISMATCH
2002×

Speed, token price and running cost are not the same discount.

Xiaomi's up to 20x faster output claim for MiMo-V2.6-Pro-UltraSpeed arrives without the test conditions behind it. Anthropic pairs a 20% token-price cut on Opus 5.501 with a separate 40% running-cost reduction claim. Treat each as a reason to test, not as an interchangeable line on your invoice.

2031: Who Keeps the Savings?

The supplied market brief reports that AMD crossed $1 trillion in market value on September 21. It puts the stock’s intraday gain as high as 10%.

Investors can value compute suppliers more highly even as developers pay less for a unit of output.

The stock move does not prove why investors bought. It does not measure installed computing capacity, either.

Lower task prices could encourage enough additional use to support higher demand for computing hardware. Savings per task need not produce lower total spending.

OpenAI’s release is a concrete case study. VentureBeat describes the new GPT-6 Luna as an enterprise workhorse model below the flagship GPT-6 Astra. The company is expanding the range of work its model family can serve.

My five-year bet is that application builders will repeatedly choose whether to keep savings or improve the product. They could also lower the customer’s bill.

The asymmetric advantage is having that choice without rebuilding the business. A product tied tightly to one provider gives up some of that freedom.

By 2031, I would want accumulated evidence about customer outcomes. Which results get approved? Which failures cause cancellations?

That knowledge can compound across model changes. Last quarter’s inference budget cannot.

Capability still determines which jobs a model can handle. Once several models clear that bar, cost per accepted task becomes a much stronger commercial headline.

Price One Publishable Clip

Build a task-margin ledger around one customer-visible result. A subtitled video highlight is a useful example because today’s digest describes Spoke as letting users highlight conversation moments and then generating subtitled video summaries.

You do not need a computer science degree to begin. A spreadsheet and a clear acceptance rule are enough for the first pass.

1. Define what the buyer receives. For a clip, specify the required format and what makes the subtitles acceptable. Keep that standard unchanged across your comparisons.

Do not pretend text-model prices cover the entire video workflow. Record editing and rendering charges separately wherever they apply.

2. Record the current delivery cost. Log model charges alongside human cleanup time. Include failed drafts even when the customer never sees them.

If you cannot observe a cost inside Spoke, mark it unknown. This exercise does not assume Spoke exposes its internal model choices.

3. Test a cheaper route you control. [Warp](https://www.warp.dev/) from today’s digest offers Workflows: reusable, parameterized terminal commands you can save and execute within Warp. Review suggested commands before running them.

Start with a text component, such as drafting a clip title or subtitle summary. Compare the candidate output against the same acceptance rule as your existing version.

4. Make the commercial decision after the test. Keep prices steady if the customer outcome holds and your offer remains valuable. Consider passing savings through when it improves a specific renewal or sale.

Avoid declaring the whole product “unlimited” from one cheap model run. Your ledger should account for the other costs first.

The same approach can track fashion concepts from [The New Black](https://thenewblack.ai/), another tool in today’s digest. Record image-generation charges separately from text-model spending.

Something will probably fail during the comparison. Keep the failed result in the ledger. That is useful evidence about the price you can safely promise.

Finish with one decision: preserve the margin, share the saving, or repair the offer. A lower model bill creates room to act. Your product determines whether you keep it.

DOJO · BUILD THIS WEEKEND

Build a task-margin ledger around one customer-visible result.

  1. Define what the buyer receives. Pick one deliverable, such as a subtitled video highlight, and write the required format plus the rule that makes it acceptable. Keep that standard identical across every comparison you run.
  2. Record the current delivery cost. Log model charges next to human cleanup time, and include failed drafts the customer never sees. Divide all workflow spending by tasks accepted as complete to get cost per accepted task.
  3. Rerun the same task on a cheaper route. Price it on GPT-6 Sol at $2 and $10 per million tokens, then on GPT-6 Luna at $0.10 and $0.50, and mark any cost you cannot observe as unknown rather than zero.
Practice: Design an AI-Assisted Workflow
THE BOTTOM LINE

Capability decides what is possible. Cost per accepted task decides what is profitable.

The September 2204 cuts hand every builder the same eight cents on the illustrative task, but not every builder gets to keep it. Products that own a workflow can spend the $80,000 monthly saving on better delivery; products that resell access will watch competitors find the same rate card. Token prices, running-cost claims and 20x02 speed claims are three different measurements, and only one of them appears on your invoice. Build the ledger around one result your customer accepts, and the next price cut becomes a choice rather than a scramble. Last quarter's inference budget cannot compound. Evidence about which results get approved can.

EDITORIAL RECEIPTKODA-20260923-BA67917B2777
As of23 September 2026MethodClaim extraction, dated-evidence review, and temporal consistency gate.CorrectionsContact the Koda desk
EVIDENCE24 CLAIMS CHECKED · 10 VERIFIED · 14 REPORTED
10 verified14 reported
  1. 01According to an Anthropic claim cited in the research brief, Opus 5.5 matches Fable 5.1 on most work.VERIFIEDTRUEBENCHMARKdigitaltrends.com
  2. 02The research brief cited in the article attributes to Xiaomi a claim of up to 20-fold faster output for MiMo-V2.6 Pro-UltraSpeed.REPORTEDMOSTLY TRUEBENCHMARKCORRECTED IN COPYmimo.xiaomi.com
  3. 03According to Xiaomi's claim cited in the research brief, MiMo-V2.6 Pro-UltraSpeed maintains the same quality as MiMo-V2.6-Pro.VERIFIEDTRUEBENCHMARKmimo.xiaomi.com
  4. 04According to Forkast, Anthropic's Opus 5.5 arrived on September 22.REPORTEDMOSTLY TRUEMODELforkast.news
  5. 05VentureBeat describes OpenAI's GPT-6 Sol as a workhorse model positioned below the Astra tier.VERIFIEDTRUEMODELventurebeat.com
  6. 06VentureBeat describes OpenAI's GPT-6 Luna as a workhorse model positioned below the Astra tier.REPORTEDMOSTLY TRUEMODELCORRECTED IN COPYventurebeat.com
  7. 07VentureBeat describes Astra as OpenAI's flagship model tier.REPORTEDMIXEDMODELCORRECTED IN COPYventurebeat.com
  8. 08Tokens are small units of text.REPORTEDMOSTLY TRUEFEATUREturingpost.com
  9. 09The AI model services discussed in the article bill for tokens.REPORTEDMIXEDFEATUREkesq.com
  10. 10Token prices, running costs, and output speed measure different aspects of AI model operation.VERIFIEDTRUEFEATUREbenchlm.ai
  11. 11An AI model wrapper is an application that adds a limited interface around another company's model.REPORTEDMOSTLY TRUEFEATURECORRECTED IN COPYstackmatix.com
  12. 12According to The Next Web, OpenAI's GPT-6 Sol is available in Codex.VERIFIEDTRUEFEATUREthenextweb.com
  13. 13According to The Next Web, OpenAI's GPT-6 Sol is available in ChatGPT Work.REPORTEDMOSTLY TRUEFEATURECORRECTED IN COPYthenextweb.com
  14. 14According to The Next Web, OpenAI's GPT-6 Luna is available in Codex.VERIFIEDTRUEFEATUREthenextweb.com
  15. 15According to The Next Web, OpenAI's GPT-6 Luna is available in ChatGPT Work.VERIFIEDTRUEFEATUREthenextweb.com
  16. 16An AI model's running cost depends on how much work the model requires to complete a task.REPORTEDMOSTLY TRUEFEATUREcapitalandcompute.net
  17. 17According to Forkast, Xiaomi released MiMo-V2.6 with open weights.VERIFIEDTRUEFEATUREcellcog.ai
  18. 18According to Forkast, Xiaomi's MiMo-V2.6 open-weight release provides a self-hosting option.REPORTEDMOSTLY TRUEFEATURECORRECTED IN COPYforkast.news
  19. 19Open model weights allow developers to run a model themselves, subject to the model's license.REPORTEDMOSTLY TRUEFEATURECORRECTED IN COPYreuters.com
  20. 20According to the digest cited in the article, Spoke creates clips from conversation highlights.REPORTEDMIXEDFEATURECORRECTED IN COPYtheneuron.ai
  21. 21Warp can help with terminal commands in a script-based workflow.REPORTEDMIXEDFEATURECORRECTED IN COPYdocs.warp.dev
  22. 22The New Black provides fashion concepts.VERIFIEDTRUEFEATUREthenewblack.ai
  23. 23Claim removed during the check; its text is not republished.REPORTEDMIXEDATTRIBUTIONCUT FROM COPYventurebeat.com
  24. 24VentureBeat reports that an OpenAI spokesperson confirmed that the new GPT-6 Sol rates are permanent.VERIFIEDTRUEATTRIBUTIONventurebeat.com

Every claim listed here was extracted from this article and checked against live sources before publication. The verdict is the checker's, not the writer's. Claims the check removed are counted but not republished.

Audit receipt KODA-20260923-BA67917B2777
Filed underMonetizationDeep Dive23 September 2026
Browse the Deep Dive archive

Get the morning Signal

181 editions so far, one a day. Unsubscribe anytime.