A 50% model price cut can raise your margin without adding a customer.
On September 22 04, its price fell to $2 per million input tokens and $10 per million output tokens. The preceding Sol rates were $4 and $20 02. Tokens are the small text units these services bill for.
That is good news if customers pay for work your product completes. It is dangerous news if they mostly pay for access to a model.
Anthropic and Xiaomi add pressure from different directions. Their claims need care: lower token prices, lower running costs, and faster output measure different things.
The useful question is who keeps the savings. We can answer it with a pricing framework and a task-level margin test.
The Savings Ownership Matrix
I think the strongest business opportunity here is to keep customer value steady while making delivery cheaper. Automatically passing every saving through can waste that opportunity.
Four figures that decide whether the cut reaches your margin.
The Savings Ownership Matrix sorts products using two questions: How much of the customer’s workflow do you own? How easily can you change models?
Workflow ownership means your product handles a meaningful piece of the job, such as a completed approval process. A blank chat box gives you much less to defend.
The matrix has four positions:
- Owners: Strong workflow value and easy model switching. Test cheaper delivery while protecting the result customers already buy.
- Dependents: Strong workflow value and difficult switching. Invest in testing and portability before assuming the announced savings are yours.
- Resellers: Weak workflow value and easy switching. Expect competitors to discover the same price cuts. Give customers a reason to stay.
- Trapped wrappers: Weak workflow value and difficult switching. Reconsider the product before spending heavily on another model migration.
An AI wrapper is typically an application that adds an interface or prompts around another company’s model. It can also add workflow logic or integrations. Being one is a starting point, but it is a fragile destination.
A cheaper GPT-6 Luna does not automatically move your business into the Owners box. You need evidence that customers value something beyond the underlying response.
Let your position guide your next investment. Owners can improve delivery. Resellers need to strengthen the offer before treating lower costs as durable profit.
The Eight-Cent Margin Test
OpenAI provides the clearest starting point. VentureBeat reports that an OpenAI spokesperson confirmed the new Sol and Luna rates are permanent.
[The Next Web](https://thenextweb.com/news/openai-gpt-6-sol-luna-api-price-cut) places both models in Codex and ChatGPT Work from September 22, 2026. The cuts reach tools people already use for work.
Sol’s input and output prices both fell 50%. Luna’s input price fell from $0.20 to $0.10 per million tokens.
Luna’s output price fell from $1.20 to $0.50, a reduction larger than half. A blanket “50% cheaper” headline hides differences in what your product actually consumes.
Consider the research brief’s illustrative task: 20,000 input tokens and 4,000 output tokens. Using OpenAI’s reported standard rates gives:
- GPT-5.6 Sol: $0.08 input plus $0.08 output, $0.16 per task
- GPT-6 Sol: $0.04 input plus $0.04 output, $0.08 per task
Assume the customer pays $1 per completed task. This is an illustration, not a reported customer contract.
Your contribution before other costs rises from $0.84 to $0.92. The gain is eight percentage points of revenue. The calculation requires neither a new customer nor a higher price.
At the brief’s illustrative volume of one million monthly tasks, model spending falls from $160,000 to $80,000. That leaves $80,000 to allocate elsewhere.
Those savings depend on unchanged task volume and quality holding up. You still need to pay migration costs.
You do not need a complicated dashboard to understand the opportunity. You need a clear invoice and a believable definition of “completed.”
Anthropic requires more caution. [Forkast](https://forkast.news/openais-gpt-6-sol-and-luna-cut-prices-50-and-the-three-tier-family-just-ended-the-coordinated-slowdown/) reports that Opus 5.5 01 arrived the same day at $4 input and $20 02 output per million tokens.
Compared with the research brief’s $5/$25 rates for Opus 5, that is a 20% 02 token-price cut. The brief separately attributes a 40% running-cost reduction claim to Anthropic.
It also cites Anthropic’s claim that Opus 5.5 01 matches Fable 5.1 on most work. It is unclear whether those comparisons use equivalent workloads.
Do not treat those claims as interchangeable discounts. Token pricing is directly visible on the rate card. Running cost depends on how much work the model needs to finish.
Xiaomi introduces another variable. The research brief attributes to Xiaomi a claim of up to 20 02× faster output speed for MiMo-V2.6-Pro-UltraSpeed at the same quality.
The supplied material does not establish the test conditions behind that claim. Treat it as a reason to test, not a promised reduction in your complete bill.
The brief also lists MiMo-V2.6-Pro at $0.435 input and $0.87 output per million tokens. At those rates, the same illustrative task costs $0.01218.
That figure covers model charges alone. It excludes the rest of the workflow.
Xiaomi’s MiMo-V2.6 open weights are also available on Hugging Face under an MIT license, according to [Forkast](https://forkast.news/xiaomis-mimo-v2-6-ships-open-weights-at-frontier-class-performance-and-the-timing-is-not-an-accident/). Open weights can let developers run the model themselves, subject to hardware and software compatibility. License restrictions also apply.
Faster generation could improve the economics of self-hosting. It does not establish a new self-hosting cost floor without hardware and operating-cost evidence.
I would measure it this way:
Cost per accepted task = all spending on the workflow ÷ tasks accepted as complete.
Include what you spend on failed attempts and human review. A result that needs repair does not have the same economics as one the customer can use immediately.
The old approach was to pick a model and price around its bill. The easier commercial test starts with one customer result and compares ways to deliver it.
Resist the temptation to name one model the cheapest for every job. Your workload decides.
Three cuts, measured in three incompatible units
The cheapest growth you will book this quarter.
At $1 per completed task, contribution before other costs rises from $0.84 to $0.92 when the model bill halves from $0.16 to $0.08. That is eight percentage points of revenue with no new customer and no price increase. It holds only if task volume and output quality stay put.
Weak workflow ownership gets undercut fast.
At the brief's listed MiMo-V2.6-Pro rates of $0.435 input and $0.87 output, the same illustrative task costs $0.01218 in model charges alone. Products that sell access rather than a completed result will meet competitors who found the same rate card. Resellers need a reason to stay before treating lower costs as durable profit.
Speed, token price and running cost are not the same discount.
Xiaomi's up to 20x faster output claim for MiMo-V2.6-Pro-UltraSpeed arrives without the test conditions behind it. Anthropic pairs a 20% token-price cut on Opus 5.5 01 with a separate 40% running-cost reduction claim. Treat each as a reason to test, not as an interchangeable line on your invoice.
2031: Who Keeps the Savings?
The supplied market brief reports that AMD crossed $1 trillion in market value on September 21. It puts the stock’s intraday gain as high as 10%.
Investors can value compute suppliers more highly even as developers pay less for a unit of output.
The stock move does not prove why investors bought. It does not measure installed computing capacity, either.
Lower task prices could encourage enough additional use to support higher demand for computing hardware. Savings per task need not produce lower total spending.
OpenAI’s release is a concrete case study. VentureBeat describes the new GPT-6 Luna as an enterprise workhorse model below the flagship GPT-6 Astra. The company is expanding the range of work its model family can serve.
My five-year bet is that application builders will repeatedly choose whether to keep savings or improve the product. They could also lower the customer’s bill.
The asymmetric advantage is having that choice without rebuilding the business. A product tied tightly to one provider gives up some of that freedom.
By 2031, I would want accumulated evidence about customer outcomes. Which results get approved? Which failures cause cancellations?
That knowledge can compound across model changes. Last quarter’s inference budget cannot.
Capability still determines which jobs a model can handle. Once several models clear that bar, cost per accepted task becomes a much stronger commercial headline.
Price One Publishable Clip
Build a task-margin ledger around one customer-visible result. A subtitled video highlight is a useful example because today’s digest describes Spoke as letting users highlight conversation moments and then generating subtitled video summaries.
You do not need a computer science degree to begin. A spreadsheet and a clear acceptance rule are enough for the first pass.
1. Define what the buyer receives. For a clip, specify the required format and what makes the subtitles acceptable. Keep that standard unchanged across your comparisons.
Do not pretend text-model prices cover the entire video workflow. Record editing and rendering charges separately wherever they apply.
2. Record the current delivery cost. Log model charges alongside human cleanup time. Include failed drafts even when the customer never sees them.
If you cannot observe a cost inside Spoke, mark it unknown. This exercise does not assume Spoke exposes its internal model choices.
3. Test a cheaper route you control. [Warp](https://www.warp.dev/) from today’s digest offers Workflows: reusable, parameterized terminal commands you can save and execute within Warp. Review suggested commands before running them.
Start with a text component, such as drafting a clip title or subtitle summary. Compare the candidate output against the same acceptance rule as your existing version.
4. Make the commercial decision after the test. Keep prices steady if the customer outcome holds and your offer remains valuable. Consider passing savings through when it improves a specific renewal or sale.
Avoid declaring the whole product “unlimited” from one cheap model run. Your ledger should account for the other costs first.
The same approach can track fashion concepts from [The New Black](https://thenewblack.ai/), another tool in today’s digest. Record image-generation charges separately from text-model spending.
Something will probably fail during the comparison. Keep the failed result in the ledger. That is useful evidence about the price you can safely promise.
Finish with one decision: preserve the margin, share the saving, or repair the offer. A lower model bill creates room to act. Your product determines whether you keep it.
Build a task-margin ledger around one customer-visible result.
- Define what the buyer receives. Pick one deliverable, such as a subtitled video highlight, and write the required format plus the rule that makes it acceptable. Keep that standard identical across every comparison you run.
- Record the current delivery cost. Log model charges next to human cleanup time, and include failed drafts the customer never sees. Divide all workflow spending by tasks accepted as complete to get cost per accepted task.
- Rerun the same task on a cheaper route. Price it on GPT-6 Sol at $2 and $10 per million tokens, then on GPT-6 Luna at $0.10 and $0.50, and mark any cost you cannot observe as unknown rather than zero.
Capability decides what is possible. Cost per accepted task decides what is profitable.
The September 22 04 cuts hand every builder the same eight cents on the illustrative task, but not every builder gets to keep it. Products that own a workflow can spend the $80,000 monthly saving on better delivery; products that resell access will watch competitors find the same rate card. Token prices, running-cost claims and 20x 02 speed claims are three different measurements, and only one of them appears on your invoice. Build the ledger around one result your customer accepts, and the next price cut becomes a choice rather than a scramble. Last quarter's inference budget cannot compound. Evidence about which results get approved can.
