OpenAI launched GPT-6 Astra, its flagship model, on September 3, 2026 02. Twenty-six days later, at DevDay on September 29, it shipped a cheaper model that ties it on a hard coding test. OpenAI says GPT-6.1 Sol matches Astra on DeepSWE v1.1 at roughly one-fifth of the cost.
The bill makes the case. Astra costs $10 09 per million input tokens and $50 10 per million output tokens. Sol costs $2 11 and $10 12. That is 80% 26 off both lines.
One admission before the numbers. Every benchmark figure below comes from OpenAI, or from outlets like DataCamp and Vellum working from OpenAI's launch material. I have not seen an independent replication yet. So read this as a strong bet you can check against your own code in a few days.
The bet: make Sol your default, and pay Astra's 5x 27 premium only on tasks where you have measured a gap worth paying for.
The Measured Gap Rule
The Measured Gap Rule fits in one line: every coding task starts on Sol and moves up to Astra only when your own logs justify it. Price sets the default. Data sets the exceptions.
What Sol-first routing costs and keeps on one benchmark
In practice, sort your work into three lanes.
Lane one is Sol only. These are tasks where a test suite is the gate: bug fixes with a failing test, new unit tests, CRUD endpoints, API glue code, docstrings. If Sol gets it wrong, the tests say so, and the retry costs a fraction of one Astra run.
Lane two is Sol first, with Astra on failure. Multi-file refactors and deep codebase investigations go here. OpenAI pitches Sol at exactly this work, calling it "built for complex refactors, deep codebase investigations, and long-running agents." Promote a task when Sol fails your verifier twice.
Lane three is Astra from the start. Database migrations and auth changes belong here. Treat payment code and anything with weak tests the same way. One patch that passes tests but quietly corrupts data can wipe out months of routine savings.
OpenAI's own safety material supports this lane. In a simulation of more than 54,000 internal Codex tasks, Astra was flagged for higher-severity misaligned behavior about half as often as the older GPT-5.6 Sol. On OpenAI's matched set for GPT-6.1 Sol the two were level: 28 severe flags for Sol, 27 for Astra. That measures behavior during autonomous runs. It says nothing direct about code correctness, and neither model has earned a long leash on risky code.
The rule is there to kill two expensive habits: paying flagship rates for boilerplate, and trusting a mid-tier model with a migration.
Where Sol Holds and Where It Slips
Think of Sol as a sharp mid-level engineer who bills at a fifth of the senior's rate. Astra is the senior you call when the mid-level gets stuck. Two questions matter: how often that happens, and what each call costs in dollars. Everything else on the launch page is noise.
Here is the scorecard from OpenAI's numbers:
- DeepSWE v1.1, long-horizon engineering in real codebases: OpenAI lists Sol at 75.2% 15 at high reasoning effort, and Astra's launch page listed 74.1% 16. Those came from separate launch runs, so treat the 1.1 28-point spread as a tie. OpenAI adds that Sol beats GPT-6 Sol's best score by 6.4 17 points.
- OSWorld 2.0 19 offline, computer use: Sol hit 71.4% at maximum effort and Astra hit 73.5%, per OpenAI.
- Terminal-Bench Science 0.1 (terminal-based science and engineering work): Sol more than doubled GPT-6 Sol's score. It averaged $5.47 01 per task at maximum reasoning effort. Astra ran $23.80 14 per task.
- Error rate: OpenAI says Sol stays within 1.9 22 points of Astra across reasoning settings.
The Terminal-Bench Science line is where the dollars get real. Divide $23.80 14 by $5.47 29 and you get 4.35, so Astra costs 4.35x more per task. That's a 77% 30 saving, a little under the 80% 26 headline. At 10,000 31 tasks, Sol runs about $54,700 and Astra about $238,000, which leaves $183,300 32 in your budget.
The retry math is the sick part. On that benchmark, Sol could attempt a task four times for $21.88 33 and still come in under one Astra run. Under the lane-two rule, the worst case is two failed Sol tries plus one Astra run: $10.94 34 plus $23.80, or $34.74. And you only pay that worst case on the tasks Sol can't finish.
DataCamp points out the catch. Sol's raw capability still trails Astra on the hardest cyber and biology evaluations. The same holds for the hardest science evaluations. If your work lives in that tail, your lane three gets bigger.
Token price and task price are different animals. OpenAI's own guidance says Astra can use substantially fewer output tokens, which lowers its cost per task. Artificial Analysis data from September 28 shows what that looks like in a real harness. Codex on Astra averaged $7.47 35 and 29.4 minutes per task, while Codex on the older GPT-6 Sol averaged $2.99 and 22.3 minutes.
So the realized gap there was about 2.5x 35, against a 5x 36 list-price gap. Those figures predate 6.1 Sol, so treat them as a hint about the pattern. Even so, the cheaper model finished faster.
One detail on the pricing card matters more than the headline, I think. Cached input on Sol costs $0.10 24 per million tokens, versus $1 37 on Astra, a 10x gap. Cached input is context the API has already seen, like the repo files an agent rereads every turn. Coding agents reread the same files constantly, so this line quietly drives a big share of the bill.
The dollars above also leave out human review. If Sol's patches take a senior engineer longer to check, that time can eat the savings, so log review minutes too.
Speed now has its own price tag. At the same DevDay, OpenAI announced an Ultrafast tier claiming up to 8x 04 faster token output in Codex and 6x 05 in the API, plus a new Pro 500 plan. Sol's fast mode is unavailable with EU data residency, so check that before you plan around it.
Where the real bill departs from list price
Four Sol attempts still undercut one Astra run.
On Terminal-Bench Science 0.1, four Sol tries cost $21.88 against $23.80 for a single Astra run. That headroom is what makes promoting a task only after two verifier failures cheap.
The realized gap can be half the list gap.
Artificial Analysis measured Codex on Astra at $7.47 per task versus $2.99 on the older GPT-6 Sol, about 2.5x against a 5x 36 list gap, because Astra can use fewer output tokens. Those figures predate 6.1 Sol, so rerun the comparison on your own code.
Sol and Astra drew about as many severe flags.
In a simulation of more than 54,000 internal Codex tasks, Astra was flagged for higher-severity misaligned behavior about half as often as the older GPT-5.6 Sol; GPT-6.1 Sol drew 28 severe flags to Astra's 27. That does not measure correctness, but it argues for keeping migrations, auth and payment code in lane three.
2031: The Router Outlives Every Model
Astra shipped September 3 at $10 09 and $50 10 per million tokens. GPT-6 Sol and Luna followed on September 22, alongside a 50% 23 price cut that Vellum reported. By September 29, Sol matched Astra's DeepSWE result at a fifth of the price.
Flagship coding performance on one benchmark reached mid-tier pricing in 26 days 38. Nobody knows whether that pace holds, and one benchmark is a thin base for a trend line, so assume something much slower. Say the price of a fixed capability only halves once a year. The $23.80 14 task then costs $11.90 39 in 2027 and about $0.74 by 2031.
That math tells you where the value compounds. The model is the part that changes every few weeks. Your eval set and your routing rules get better each time you swap models. Your cost logs grow more useful too. A team with a router moves to the next cheap model by editing one line of config.
The risk is lopsided. On a lane-one task, a wrong Sol answer costs one $5.47 40 retry. Hardcoding Astra everywhere means paying 4.35x 29 on every task, forever. Sol-first only really burns you on lane-three work, which is why that lane exists.
My read on OpenAI's packaging is that the flagship premium is moving. Ultrafast and Pro 500 suggest OpenAI expects to sell speed and headroom once mid-tier models catch up on raw capability. By 2031, I expect premium tiers to compete mostly on latency and tail reliability. If that's right, the "which model is smartest" question matters less every quarter.
Price Your Coding Agent Runs by Friday
You do not need a CS degree for this. You need a spreadsheet, an API key and a short script. Plan on one afternoon for setup and one for runs.
First, pull 50 merged pull requests from the last two months where you already know the right answer. Tag each one with its lane using the Measured Gap Rule. Aim for a mix that looks like your real week.
Second, run every task through both gpt-6.1-sol and gpt-6-astra with an identical harness and prompt. A harness is the script that feeds the task to the model and runs your tests on the result. Log input tokens, cached tokens, output tokens, attempts, wall-clock time, review minutes and pass or fail.
Third, compute cost per accepted result. Take total spend for each model and divide by the number of tasks that passed. That single number beats any leaderboard, because it comes from your code.
Fourth, write the router. It can be twenty lines. Lane-three tags go straight to Astra. Everything else starts on Sol, and a task that fails your tests twice gets promoted.
Things will break on the first pass. Your first run will hit flaky tests and a harness bug or two. That is normal, and each fix makes the next model swap faster. Rerun the table when the Ultrafast tier lands, since speed is now a line item too.
Then post your cost-per-result table, even if it is ugly. OpenAI's benchmarks describe OpenAI's test setup. Your 50 tasks are the first real evidence about your codebase, and they cost you a couple of afternoons.
Price your own coding agent runs before trusting a leaderboard
- Pull 50 known-good tasks. Take 50 merged pull requests from the last two months and tag each one lane one, two or three using the Measured Gap Rule.
- Run both models through one harness. Send every task to gpt-6.1-sol and gpt-6-astra with the same prompt, and log input, cached and output tokens, attempts, wall-clock time, review minutes and pass or fail.
- Compute cost per accepted result, then route. Divide each model's total spend by the tasks that passed, then write a twenty-line router that sends lane-three tags to Astra and promotes any Sol task that fails your tests twice.
Make Sol the default and make Astra earn its premium.
Every benchmark here comes from OpenAI's own launch material, so treat the tie as a bet to verify, not a fact. The price gap is large enough that checking costs little: two afternoons, 50 tasks and a cost-per-result table. Keep migrations, auth and payment code on Astra, route everything else through Sol, and let your logs decide when to promote a task. The model will change within weeks, but your eval set and your router will carry over to each new one.
