K Koda Intelligence
DEEP DIVE DEEP DIVE № 212 · 30 September 2026DOCKODA-20260930-A33A351785F3sha-256 of date + article + 24 checked + 16 computed

GPT-6.1 Sol ties Astra on coding at a fifth of the price

At DevDay on September 29, OpenAI shipped GPT-6.1 Sol, listed at 75.2% 15 on DeepSWE v1.1 against Astra's 74.1% 16. It costs $2 11 and $10 09 per million tokens, 80% 26 below Astra on both lines. On OSWorld 2.0 19 it scores 71.4% to Astra's 73.5%. The practical move is to default to Sol and pay Astra's premium only where your own logs show a gap.

7 MIN READ · BY THE KODA EDITORIAL TEAM · TOOLS · MODEL ROUTING
SOL DEEPSWE V1.175.2%VERIFIED CLAIM 15↑ 6.4 PTS
ASTRA DEEPSWE V1.174.1%REPORTED CLAIM 16OpenAI
LIST PRICE CUT80%COMPUTED CLAIM 26INPUT AND OUTPUT
SOL DEEPSWE V1.175.2%↑ 6.4 PTS ASTRA DEEPSWE V1.174.1%OpenAI LIST PRICE CUT80%INPUT AND OUTPUT SOL COST PER TASK$5.47TERMINAL-BENCH SCIENCE ASTRA COST PER TASK$23.80TERMINAL-BENCH SCIENCE SOL CACHED INPUT$0.10PER MILLION TOKENS DAYS TO MATCH26SEP 3 TO SEP 29

OpenAI launched GPT-6 Astra, its flagship model, on September 3, 2026 02. Twenty-six days later, at DevDay on September 29, it shipped a cheaper model that ties it on a hard coding test. OpenAI says GPT-6.1 Sol matches Astra on DeepSWE v1.1 at roughly one-fifth of the cost.

The bill makes the case. Astra costs $10 09 per million input tokens and $50 10 per million output tokens. Sol costs $2 11 and $10 12. That is 80% 26 off both lines.

One admission before the numbers. Every benchmark figure below comes from OpenAI, or from outlets like DataCamp and Vellum working from OpenAI's launch material. I have not seen an independent replication yet. So read this as a strong bet you can check against your own code in a few days.

The bet: make Sol your default, and pay Astra's 5x 27 premium only on tasks where you have measured a gap worth paying for.

The Measured Gap Rule

The Measured Gap Rule fits in one line: every coding task starts on Sol and moves up to Astra only when your own logs justify it. Price sets the default. Data sets the exceptions.

COST PER TASK MATH · SEPTEMBER 2026OpenAI · TERMINAL-BENCH SCIENCE 0.1BASE: 24 CHECKED + 16 COMPUTED, 4 SHOWN

What Sol-first routing costs and keeps on one benchmark

Sol per task OpenAI · max reasoning effort REPORTED CLAIM 01
$5.47
Astra per task OpenAI · same benchmark REPORTED CLAIM 14
$23.80
Lane-two worst case Two failed Sol tries plus one Astra run COMPUTED CLAIM 34
$34.74
Kept at 10,000 tasks Sol vs Astra · $54,700 vs $238,000 COMPUTED CLAIM 32
$183,300

In practice, sort your work into three lanes.

Lane one is Sol only. These are tasks where a test suite is the gate: bug fixes with a failing test, new unit tests, CRUD endpoints, API glue code, docstrings. If Sol gets it wrong, the tests say so, and the retry costs a fraction of one Astra run.

Lane two is Sol first, with Astra on failure. Multi-file refactors and deep codebase investigations go here. OpenAI pitches Sol at exactly this work, calling it "built for complex refactors, deep codebase investigations, and long-running agents." Promote a task when Sol fails your verifier twice.

Lane three is Astra from the start. Database migrations and auth changes belong here. Treat payment code and anything with weak tests the same way. One patch that passes tests but quietly corrupts data can wipe out months of routine savings.

OpenAI's own safety material supports this lane. In a simulation of more than 54,000 internal Codex tasks, Astra was flagged for higher-severity misaligned behavior about half as often as the older GPT-5.6 Sol. On OpenAI's matched set for GPT-6.1 Sol the two were level: 28 severe flags for Sol, 27 for Astra. That measures behavior during autonomous runs. It says nothing direct about code correctness, and neither model has earned a long leash on risky code.

The rule is there to kill two expensive habits: paying flagship rates for boilerplate, and trusting a mid-tier model with a migration.

Where Sol Holds and Where It Slips

Think of Sol as a sharp mid-level engineer who bills at a fifth of the senior's rate. Astra is the senior you call when the mid-level gets stuck. Two questions matter: how often that happens, and what each call costs in dollars. Everything else on the launch page is noise.

Here is the scorecard from OpenAI's numbers:

  • DeepSWE v1.1, long-horizon engineering in real codebases: OpenAI lists Sol at 75.2% 15 at high reasoning effort, and Astra's launch page listed 74.1% 16. Those came from separate launch runs, so treat the 1.1 28-point spread as a tie. OpenAI adds that Sol beats GPT-6 Sol's best score by 6.4 17 points.
  • OSWorld 2.0 19 offline, computer use: Sol hit 71.4% at maximum effort and Astra hit 73.5%, per OpenAI.
  • Terminal-Bench Science 0.1 (terminal-based science and engineering work): Sol more than doubled GPT-6 Sol's score. It averaged $5.47 01 per task at maximum reasoning effort. Astra ran $23.80 14 per task.
  • Error rate: OpenAI says Sol stays within 1.9 22 points of Astra across reasoning settings.

The Terminal-Bench Science line is where the dollars get real. Divide $23.80 14 by $5.47 29 and you get 4.35, so Astra costs 4.35x more per task. That's a 77% 30 saving, a little under the 80% 26 headline. At 10,000 31 tasks, Sol runs about $54,700 and Astra about $238,000, which leaves $183,300 32 in your budget.

The retry math is the sick part. On that benchmark, Sol could attempt a task four times for $21.88 33 and still come in under one Astra run. Under the lane-two rule, the worst case is two failed Sol tries plus one Astra run: $10.94 34 plus $23.80, or $34.74. And you only pay that worst case on the tasks Sol can't finish.

DataCamp points out the catch. Sol's raw capability still trails Astra on the hardest cyber and biology evaluations. The same holds for the hardest science evaluations. If your work lives in that tail, your lane three gets bigger.

Token price and task price are different animals. OpenAI's own guidance says Astra can use substantially fewer output tokens, which lowers its cost per task. Artificial Analysis data from September 28 shows what that looks like in a real harness. Codex on Astra averaged $7.47 35 and 29.4 minutes per task, while Codex on the older GPT-6 Sol averaged $2.99 and 22.3 minutes.

So the realized gap there was about 2.5x 35, against a 5x 36 list-price gap. Those figures predate 6.1 Sol, so treat them as a hint about the pattern. Even so, the cheaper model finished faster.

One detail on the pricing card matters more than the headline, I think. Cached input on Sol costs $0.10 24 per million tokens, versus $1 37 on Astra, a 10x gap. Cached input is context the API has already seen, like the repo files an agent rereads every turn. Coding agents reread the same files constantly, so this line quietly drives a big share of the bill.

The dollars above also leave out human review. If Sol's patches take a senior engineer longer to check, that time can eat the savings, so log review minutes too.

Speed now has its own price tag. At the same DevDay, OpenAI announced an Ultrafast tier claiming up to 8x 04 faster token output in Codex and 6x 05 in the API, plus a new Pro 500 plan. Sol's fast mode is unavailable with EU data residency, so check that before you plan around it.

Where the real bill departs from list price

RETRY HEADROOM
$21.88 33

Four Sol attempts still undercut one Astra run.

On Terminal-Bench Science 0.1, four Sol tries cost $21.88 against $23.80 for a single Astra run. That headroom is what makes promoting a task only after two verifier failures cheap.

TOKEN EFFICIENCY
2.5x 35

The realized gap can be half the list gap.

Artificial Analysis measured Codex on Astra at $7.47 per task versus $2.99 on the older GPT-6 Sol, about 2.5x against a 5x 36 list gap, because Astra can use fewer output tokens. Those figures predate 6.1 Sol, so rerun the comparison on your own code.

AUTONOMY RISK
54,000

Sol and Astra drew about as many severe flags.

In a simulation of more than 54,000 internal Codex tasks, Astra was flagged for higher-severity misaligned behavior about half as often as the older GPT-5.6 Sol; GPT-6.1 Sol drew 28 severe flags to Astra's 27. That does not measure correctness, but it argues for keeping migrations, auth and payment code in lane three.

2031: The Router Outlives Every Model

Astra shipped September 3 at $10 09 and $50 10 per million tokens. GPT-6 Sol and Luna followed on September 22, alongside a 50% 23 price cut that Vellum reported. By September 29, Sol matched Astra's DeepSWE result at a fifth of the price.

Flagship coding performance on one benchmark reached mid-tier pricing in 26 days 38. Nobody knows whether that pace holds, and one benchmark is a thin base for a trend line, so assume something much slower. Say the price of a fixed capability only halves once a year. The $23.80 14 task then costs $11.90 39 in 2027 and about $0.74 by 2031.

That math tells you where the value compounds. The model is the part that changes every few weeks. Your eval set and your routing rules get better each time you swap models. Your cost logs grow more useful too. A team with a router moves to the next cheap model by editing one line of config.

The risk is lopsided. On a lane-one task, a wrong Sol answer costs one $5.47 40 retry. Hardcoding Astra everywhere means paying 4.35x 29 on every task, forever. Sol-first only really burns you on lane-three work, which is why that lane exists.

My read on OpenAI's packaging is that the flagship premium is moving. Ultrafast and Pro 500 suggest OpenAI expects to sell speed and headroom once mid-tier models catch up on raw capability. By 2031, I expect premium tiers to compete mostly on latency and tail reliability. If that's right, the "which model is smartest" question matters less every quarter.

Price Your Coding Agent Runs by Friday

You do not need a CS degree for this. You need a spreadsheet, an API key and a short script. Plan on one afternoon for setup and one for runs.

First, pull 50 merged pull requests from the last two months where you already know the right answer. Tag each one with its lane using the Measured Gap Rule. Aim for a mix that looks like your real week.

Second, run every task through both gpt-6.1-sol and gpt-6-astra with an identical harness and prompt. A harness is the script that feeds the task to the model and runs your tests on the result. Log input tokens, cached tokens, output tokens, attempts, wall-clock time, review minutes and pass or fail.

Third, compute cost per accepted result. Take total spend for each model and divide by the number of tasks that passed. That single number beats any leaderboard, because it comes from your code.

Fourth, write the router. It can be twenty lines. Lane-three tags go straight to Astra. Everything else starts on Sol, and a task that fails your tests twice gets promoted.

Things will break on the first pass. Your first run will hit flaky tests and a harness bug or two. That is normal, and each fix makes the next model swap faster. Rerun the table when the Ultrafast tier lands, since speed is now a line item too.

Then post your cost-per-result table, even if it is ugly. OpenAI's benchmarks describe OpenAI's test setup. Your 50 tasks are the first real evidence about your codebase, and they cost you a couple of afternoons.

DOJO · BUILD THIS WEEKEND

Price your own coding agent runs before trusting a leaderboard

  1. Pull 50 known-good tasks. Take 50 merged pull requests from the last two months and tag each one lane one, two or three using the Measured Gap Rule.
  2. Run both models through one harness. Send every task to gpt-6.1-sol and gpt-6-astra with the same prompt, and log input, cached and output tokens, attempts, wall-clock time, review minutes and pass or fail.
  3. Compute cost per accepted result, then route. Divide each model's total spend by the tasks that passed, then write a twenty-line router that sends lane-three tags to Astra and promotes any Sol task that fails your tests twice.
Practice: Design an AI-Assisted Workflow
THE BOTTOM LINE

Make Sol the default and make Astra earn its premium.

Every benchmark here comes from OpenAI's own launch material, so treat the tie as a bet to verify, not a fact. The price gap is large enough that checking costs little: two afternoons, 50 tasks and a cost-per-result table. Keep migrations, auth and payment code on Astra, route everything else through Sol, and let your logs decide when to promote a task. The model will change within weeks, but your eval set and your router will carry over to each new one.

LISTEN · AUDIO BRIEFINGThe conversation · ~20 min
WATCH · VISUAL NARRATIVEAnimated breakdown · ~6 min
PLAY · YOUTUBE
EDITORIAL RECEIPTKODA-20260930-A33A351785F3
As of30 September 2026MethodClaim extraction, dated-evidence review, and temporal consistency gate.CorrectionsContact the Koda desk
EVIDENCE24 CHECKED + 16 COMPUTED · 12 VERIFIED · 11 REPORTED · 1 FAILED
12 verified11 reported1 failed16 computed
  1. 01- Terminal-Bench Science 0.1, terminal-based science and engineering work: Sol more than doubled GPT-6 Sol's score at $5.47 per task on average at maximum reasoning effort.REPORTEDUNVERIFIABLEBENCHMARKCORRECTED IN COPYstartupfortune.com
  2. 02OpenAI launched GPT-6 Astra, its flagship model, on September 3, 2026.REPORTEDMOSTLY TRUEMODELopenai.com
  3. 03OpenAI shipped GPT-6.1 Sol at its DevDay event on September 29, 2026.REPORTEDMOSTLY TRUEMODELopenai.com
  4. 04At DevDay on September 29, 2026, OpenAI announced an Ultrafast tier claiming up to 8x faster token output in Codex.VERIFIEDTRUEFEATUREopenai.com
  5. 05At DevDay on September 29, 2026, OpenAI announced an Ultrafast tier claiming up to 6x faster token output in the API.REPORTEDMOSTLY TRUEFEATUREopenai.com
  6. 06At DevDay on September 29, 2026, OpenAI announced a new Pro 500 plan.VERIFIEDTRUEFEATUREopenai.com
  7. 07OpenAI says GPT-6.1 Sol matches GPT-6 Astra on DeepSWE v1.1 at roughly one-fifth of the cost.VERIFIEDTRUEATTRIBUTIONopenai.com
  8. 08The GPT-6.1 Sol and GPT-6 Astra benchmark figures come from OpenAI, or from outlets such as DataCamp and Vellum working from OpenAI's launch material.REPORTEDMOSTLY TRUEATTRIBUTIONopenai.com
  9. 09GPT-6 Astra costs $10 per million input tokens.VERIFIEDTRUEPRICEhelp.openai.com
  10. 10GPT-6 Astra costs $50 per million output tokens.VERIFIEDTRUEPRICEopenai.com
  11. 11GPT-6.1 Sol costs $2 per million input tokens.VERIFIEDTRUEPRICEopenai.com
  12. 12GPT-6.1 Sol costs $10 per million output tokens.VERIFIEDTRUEPRICEteamday.ai
  13. 13- Terminal-Bench Science 0.1, terminal-based science and engineering work: Sol more than doubled GPT-6 Sol's score at $5.47 per task on average at maximum reasoning effort.REPORTEDMOSTLY TRUESTATCORRECTED IN COPYopenai.com
  14. 14GPT-6 Astra cost $23.80 per task on Terminal-Bench Science 0.1.REPORTEDMOSTLY TRUESTATopenai.com
  15. 15OpenAI lists GPT-6.1 Sol at 75.2% on DeepSWE v1.1 at high reasoning effort.VERIFIEDTRUEBENCHMARKopenai.com
  16. 16GPT-6 Astra's launch page listed a DeepSWE v1.1 score of 74.1%.REPORTEDMOSTLY TRUEBENCHMARKopenai.com
  17. 17OpenAI says GPT-6.1 Sol beats GPT-6 Sol's best DeepSWE v1.1 score by 6.4 points.VERIFIEDTRUEBENCHMARKopenai.com
  18. 18GPT-6.1 Sol scored 71.4% on OSWorld 2.0 offline (computer use) at maximum effort.VERIFIEDTRUEBENCHMARKopenai.com
  19. 19- OSWorld 2.0 offline, computer use: Sol hit 71.4% at maximum effort and Astra hit 72.6%, per OpenAI.FAILEDMOSTLY FALSEBENCHMARKCORRECTED IN COPYdevelopersdigest.tech
  20. 20OpenAI released GPT-6 Sol and GPT-6 Luna on September 22, 2026.VERIFIEDTRUEMODELopenai.com
  21. 21Claim removed during the check; its text is not republished.REPORTEDMIXEDATTRIBUTIONCUT FROM COPYopenai.com
  22. 22OpenAI says GPT-6.1 Sol's error rate stays within 1.9 points of GPT-6 Astra's across reasoning settings.REPORTEDMOSTLY TRUEATTRIBUTIONopenai.com
  23. 23Vellum reported a 50% price cut from OpenAI alongside the September 22, 2026 release of GPT-6 Sol and GPT-6 Luna.REPORTEDMOSTLY TRUEATTRIBUTIONwebsite.vellum.ai
  24. 24Cached input on GPT-6.1 Sol costs $0.10 per million tokens.VERIFIEDTRUEPRICEopenai.com
  25. 25GPT-6.1 Sol shipped 26 days after GPT-6 Astra (September 3 to September 29, 2026).COMPUTEDCOMPUTED
  26. 26GPT-6.1 Sol's pricing is 80% lower than GPT-6 Astra's on both input and output tokens.COMPUTEDCOMPUTED
  27. 27GPT-6 Astra carries a 5x price premium over GPT-6.1 Sol.COMPUTEDCOMPUTED
  28. 28GPT-6.1 Sol's DeepSWE v1.1 score is 1.1 points higher than GPT-6 Astra's (75.2% vs 74.1%).COMPUTEDCOMPUTED
  29. 29On Terminal-Bench Science 0.1, GPT-6 Astra costs 4.35x more per task than GPT-6.1 Sol ($23.80 divided by $5.47).COMPUTEDCOMPUTED
  30. 30On Terminal-Bench Science 0.1, GPT-6.1 Sol's per-task cost is a 77% saving versus GPT-6 Astra.COMPUTEDCOMPUTED
  31. 31At 10,000 Terminal-Bench Science 0.1 tasks, GPT-6.1 Sol would cost about $54,700 and GPT-6 Astra about $238,000.COMPUTEDCOMPUTED
  32. 32Running 10,000 Terminal-Bench Science 0.1 tasks on GPT-6.1 Sol instead of GPT-6 Astra saves $183,300.COMPUTEDCOMPUTED
  33. 33On Terminal-Bench Science 0.1, four GPT-6.1 Sol attempts cost $21.88, less than one GPT-6 Astra run at $23.80.COMPUTEDCOMPUTED
  34. 34The worst case of two failed GPT-6.1 Sol tries plus one GPT-6 Astra run on Terminal-Bench Science 0.1 costs $34.74 ($10.94 plus $23.80).COMPUTEDCOMPUTED
  35. 35The realized per-task cost gap between Codex on GPT-6 Astra ($7.47) and Codex on GPT-6 Sol ($2.99) was about 2.5x.COMPUTEDCOMPUTED
  36. 36The list-price gap between GPT-6 Astra and GPT-6.1 Sol is 5x.COMPUTEDCOMPUTED
  37. 37GPT-6 Astra's cached input price is 10x GPT-6.1 Sol's ($1 vs $0.10 per million tokens).COMPUTEDCOMPUTED
  38. 38Flagship coding performance on DeepSWE v1.1 reached mid-tier pricing in 26 days (GPT-6 Astra on September 3 to GPT-6.1 Sol on September 29, 2026).COMPUTEDCOMPUTED
  39. 39If the price of a fixed capability halves once a year, the $23.80 GPT-6 Astra Terminal-Bench Science task would cost $11.90 in 2027 and about $0.74 by 2031.COMPUTEDCOMPUTED
  40. 40A wrong GPT-6.1 Sol answer on a lane-one task costs one $5.47 retry.COMPUTEDCOMPUTED

Every claim listed here was extracted from this article and checked against live sources before publication. The verdict is the checker's, not the writer's. Claims the check removed are counted but not republished.

Audit receipt KODA-20260930-A33A351785F3
Filed underToolsDeep Dive30 September 2026
Browse the Deep Dive archive

Get the morning Signal

187 editions so far, one a day. Unsubscribe anytime.