K Koda Intelligence
DEEP DIVE DEEP DIVE № 227 · 09 October 2026DOCKODA-20261009-36460207D9C6sha-256 of date + article + 24 checked + 13 computed

Same price, 23.5 points apart on agent work

Anthropic reports Claude Haiku 5.5 scored 72.4% 25 on OSWorld 2.1, against 48.9% for GPT-6 Luna, at the same sticker price. Luna reaches ChatGPT Free and Go users from October 9, so the budget tier is where most people meet AI. Every head-to-head figure comes from Anthropic's own harness. We show how to price a finished task and test both models on your own work.

6 MIN READ · BY THE KODA EDITORIAL TEAM · TOOLS · AGENT ECONOMICS
HAIKU 5.5 OSWORLD72.4%VERIFIED CLAIM 01↑ 23.5 PTS VS GPT-6 LUNA
GPT-6 LUNA OSWORLD48.9%VERIFIED CLAIM 02Anthropic
HAIKU FRONTIERCODE46.4%VERIFIED CLAIM 16↑ 4 PTS VS GPT-6 LUNA
HAIKU 5.5 OSWORLD72.4%↑ 23.5 PTS VS GPT-6 LUNA GPT-6 LUNA OSWORLD48.9%Anthropic HAIKU FRONTIERCODE46.4%↑ 4 PTS VS GPT-6 LUNA HAIKU INPUT PRICE$0.10FROM $1.00 ON HAIKU 4.5 AVG RUN COST CUT75%Anthropic SONNET 5.5 OSWORLD83.9%AIIDELIST.COM LINEUP REFRESH15 DAYSPROJEDEFTERI.COM

Anthropic's cheapest new model just beat OpenAI's cheapest new model by 23.5 25 points on a test of operating a computer. The two cost the same. Anthropic reports Claude Haiku 5.5 scored 72.4% 01 on the OSWorld 2.1 offline subset. GPT-6 Luna scored 48.9% 02.

Anthropic's release page lists Haiku 5.5 at $0.10 11 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens, and $0.50 and $2.50 above that. moelueker.com's launch review notes that, up to 100K tokens, that is exactly GPT-6 Luna's sticker price. Same price, very different results on agent work.

The cheap tier is now the front line. OpenAI is rolling Luna out to ChatGPT Free and Go users from October 9, while paid tiers get GPT-6 Sol. For millions of people, the budget model is the only one they will ever touch.

One headline from this week needs correcting first. We could not find a Microsoft announcement or Azure pricing page confirming that, so treat the Microsoft 75% 09 figure as unverified. The 75% cut we can source belongs to Anthropic, and it tells a more useful story.

Same sticker price, different finished jobs

The sticker gap is the distance between what a model costs per token and what it costs per job that actually finishes. This week's numbers show it twice.

COST PER FINISHED TASK · OCTOBER 2026Anthropic · PPC.LAND · KODA ANALYSISBASE: 24 CHECKED + 13 COMPUTED, 4 SHOWN

What a finished agent task actually costs on the cheap tier

Haiku 5.5 input price Anthropic · per million tokens, prompts up to 100K REPORTED CLAIM 11
$0.10
Long-prompt threshold Anthropic · above this, rates rise five times REPORTED CLAIM 11
100K
Haiku 5.5 per finished task Our estimate · $0.20 attempt × 1.38 attempts COMPUTED CLAIM 33
$0.28
GPT-6 Luna per finished task Our estimate · $0.20 attempt × 2.04 attempts COMPUTED CLAIM 34
$0.41

The first case is Haiku 5.5 against Luna. On Anthropic's OSWorld figures, Haiku needs about 1.38 26 attempts per success (1 divided by 0.724). Luna needs about 2.04 27 attempts (1 divided by 0.489). Same token price, and Luna burns roughly 48% 28 more attempts for every finished task.

The second case is Anthropic's own price cut. ppc.land reported that the rate card fell from $1.00 13 and $5.00 on Haiku 4.5 to $0.10 and $0.50 on Haiku 5.5 for prompts up to 100,000 tokens, with larger prompts costing $0.50 and $2.50. Anthropic's release page says the model costs "around 75% 09 less to run" on average.

So where did the missing 15 points go? Projedefteri.com has the answer. Haiku 5.5 ships with a new tokenizer, the piece that chops text into billable units. It also has a long-prompt tier: any request over 100,000 tokens 14 costs five times more, at $0.50 in and $2.50 out.

Every agent bill comes down to three things. First, the price per attempt: tokens in and out, at whichever tier the prompt lands in. Second, the success rate, meaning the share of attempts that finish the job. Divide the attempt price by it. Third, cleanup: the retries and human fixes on the runs that failed.

Brand names show up in none of them.

Reading Anthropic's table like an ops lead

Picture paying a courier per trip. The logo on the van doesn't matter. You care about parcels delivered per trip, and how many came back damaged.

Anthropic's comparison table gives you the delivery rates, and the margins swing hard by job type. Haiku 5.5 leads Luna 72.4% 25 to 48.9% on OSWorld 2.1. On FrontierCode 1.1 29 the lead shrinks to 4 points, 46.4% against 42.4%.

That spread is the whole lesson. A 23.5-point 25 lead on clicking through desktop apps says little about a coding task where the gap is 4. Pick the benchmark that looks most like your job and ignore the rest.

The bigger sibling still matters. Aiidelist.com, reporting Anthropic's numbers, puts Claude Sonnet 5.5 at 83.9% 04 on OSWorld's offline subset (Anthropic's headline score was 80.1%) and 70.6% on Terminal-Bench. Haiku's 39.2% 30 on the terminal test is a bit over half of Sonnet's. Projedefteri.com's advice follows from that: use Haiku for summaries and browser work, and keep it out of the lead seat on complex agentic coding.

Anthropic's own launch page says the same thing more gently. It pitches Haiku 5.5 as a subagent, a model handed one slice of a bigger job, working under Opus 5.5 or Sonnet 5.5. Anthropic also halved Sonnet 5.5's cache-read price, which it says makes Sonnet around 20% 21 cheaper on most agentic work. So the escalation model got cheaper on the same day the cheap model got better.

Take a job that sends 1 million 31 input tokens and gets back 200,000 output tokens. Under 100,000 tokens per request, that's $0.10 plus $0.10, so $0.20.

Push the same volume through prompts over 100,000 tokens 32 and it costs $0.50 plus $0.50, so $1.00. How fast your agent's context grows decides which tier you live in.

Retries widen the gap. Assume both models spend $0.20 33 per attempt. Haiku costs about $0.28 per finished task ($0.20 times 1.38). Luna costs about $0.41 34 ($0.20 times 2.04). The equal-tokens-per-attempt assumption is ours, and real runs will vary.

Every Haiku-versus-Luna number above comes from Anthropic's own evaluation setup, and that is the weak point. BenchLM.ai's comparison page, updated October 7, says the public evidence has no benchmark result shared by both models. In BenchLM's words, that evidence "does not support a quality verdict."

Nobody outside Anthropic has shown whether the 23.5-point 25 gap survives a different harness with Luna run at its best settings. We think the table justifies a test. It doesn't justify a migration. The test is cheap, and the next section shows why it keeps paying.

Hidden costs behind identical sticker prices

UNVERIFIED LEAD
23.5 25

Anthropic's OSWorld lead lacks outside confirmation.

Every Haiku-versus-Luna number comes from Anthropic's own evaluation setup. BenchLM.ai found no benchmark result shared by both models, so the gap justifies a test and still needs an independent harness.

LONG-PROMPT TIER
100K 11

Growing context quietly multiplies the bill.

Haiku 5.5 requests over 100,000 tokens 14 cost $0.50 in and $2.50 out. The same 1 million 32 input and 200,000 output tokens cost $0.20 31 below the line and $1.00 above it.

FAILURE COSTS
2031

Retries and review become most of the bill.

If small-model running costs keep falling about 75% 36 per generation, a $1.00 job drops to about $0.016 by 2031. Failed runs and human fixes do not shrink at that pace, and benchmarks like IFBench are already being retired.

2031. Count the release dates. Projedefteri.com lists Opus 5.5 on September 22, Sonnet 5.5 on September 28 and Haiku 5.5 on October 7. Anthropic refreshed its entire lineup in 15 days 35.

Measurement moved just as fast. On June 15, 2026 24, Artificial Analysis announced version 4.1 of its Intelligence Index. It reweighted the index toward agentic tasks and added three new metrics, led by Cost per Task. The other two are Time per Task and Tokens per Task. Four months later, moelueker.com's Haiku launch cover charts OSWorld score against cost per task. The scoreboard is moving toward the finished job.

The napkin math goes like this. Anthropic says one small-model generation cut average running cost by about 75% 36. Suppose that happens three more times by 2031. A $1.00 job becomes $0.25, then about $0.06, then about $0.016.

Straight lines break, so hold that loosely. The direction is still right. As token spend falls toward pennies, failure costs don't fall with it. Retries and human review end up as most of the bill.

Benchmarks age too. In the same June update, Artificial Analysis dropped IFBench from its index because frontier models had saturated it. Sonnet 5.5 already sits at 83.9% 04 on the OSWorld subset. By 2031, the test you trust this week may be retired.

Put the two trends together and the bet is lopsided. A private test set of your own tasks costs a few dollars of tokens and a day of setup. Choosing by brand costs you on every run, every month, across every launch you skip evaluating. And a test set you keep rerunning compounds: each new model takes an hour to judge, while competitors are still reading launch posts.

Score twenty real tasks before Friday

The goal is one spreadsheet that tells you which model to send each job to, with the cost of a finished task beside each name.

Step one: define done. Pick one job your agent runs daily, such as filing a support ticket or filling a web form. Write one sentence that says what a finished run looks like. Pull 20 real examples, and make sure at least 5 are ones that went wrong before.

Step two: run the bake-off. Send all 20 through Haiku 5.5 (API name claude-haiku-5-5), GPT-6 Luna and whatever model you use today. Keep the prompt, tools and stop rule identical across all three. Log input tokens, output tokens, pass or fail, retries and minutes a person spent fixing the output.

Step three: price each attempt. Divide input tokens by 1,000,000 and multiply by the input price. Do the same for output. Add the two to get cost per attempt.

Step four: price each finished task. Divide cost per attempt by the pass rate. Pass rate means the share of runs that met your one-sentence definition of done. Then add the human minutes, multiplied by what an hour of that person costs you.

Step five: flag the expensive tier. Mark every Haiku request that crossed 100,000 tokens 14. Those cost five times more. If your agent keeps crossing the line, trim its context before you blame the model.

Step six: write the routing rule. Send the job to the cheapest model whose pass rate clears your bar. Anything that fails the check escalates to Sonnet 5.5. Score the cost of that escalation path as well.

Expect things to break. Browser sessions will log out, tool calls will time out and some runs will hang. Count every one of them as a fail, because production will. Then put a reminder on the sheet: rerun all 20 tasks the day any vendor changes a price or ships a model. On this year's evidence, that could mean every few weeks.

DOJO · BUILD THIS WEEKEND

Score twenty real tasks and write a routing rule

  1. Define done and pull 20 examples. Pick one daily agent job, write a one-sentence finish line, and include at least 5 runs that failed before.
  2. Run an identical bake-off. Send all 20 through Haiku 5.5, GPT-6 Luna and your current model with the same prompt, tools and stop rule, logging tokens, pass or fail, retries and human fix minutes.
  3. Price the finished task. Divide cost per attempt by pass rate, add human minutes at their hourly cost, and flag every Haiku request over 100,000 31 tokens before routing jobs to the cheapest model that clears your bar.
Practice: Design an AI-Assisted Workflow
THE BOTTOM LINE

Price the finished job, then pick the model.

Haiku 5.5 and GPT-6 Luna share a sticker price and land far apart on Anthropic's agent tests. Those results are self-reported, so we treat them as a reason to run a test of our own. Twenty real tasks, a spreadsheet and a few dollars of tokens tell you which model finishes your work for less. Rerun it every time a vendor ships a model or changes a price, which this year means every few weeks.

LISTEN · AUDIO BRIEFINGThe conversation · ~12 min
WATCH · VISUAL NARRATIVEAnimated breakdown · ~6 min
PLAY · YOUTUBE
EDITORIAL RECEIPTKODA-20261009-36460207D9C6
As of09 October 2026MethodClaim extraction, dated-evidence review, and temporal consistency gate.CorrectionsContact the Koda desk
EVIDENCE24 CHECKED + 13 COMPUTED · 10 VERIFIED · 12 REPORTED · 2 FAILED
10 verified12 reported2 failed13 computed
  1. 01Anthropic reports that Claude Haiku 5.5 scored 72.4% on the OSWorld 2.1 offline subset.VERIFIEDTRUEBENCHMARKenterpriseai.economictimes.indiatimes.com
  2. 02GPT-6 Luna scored 48.9% on the OSWorld 2.1 offline subset, according to Anthropic's figures.VERIFIEDTRUEBENCHMARKanthropic.com
  3. 03Claude Haiku 5.5 scored 39.2% on Terminal-Bench 4.0, according to Anthropic's comparison table.VERIFIEDTRUEBENCHMARKanthropic.com
  4. 04Aiidelist.com, reporting Anthropic's numbers, puts Claude Sonnet 5.5 at 83.9% on OSWorld's offline subset (Anthropic's headline score was 80.1%) and 70.6% on Terminal-Bench.REPORTEDMIXEDBENCHMARKCORRECTED IN COPYanthropic.com
  5. 05Claude Haiku 5.5 is Anthropic's cheapest new model, and GPT-6 Luna is OpenAI's cheapest new model.REPORTEDMIXEDMODELanthropic.com
  6. 06OpenAI is rolling GPT-6 Luna out to ChatGPT Free and Go users from October 9.VERIFIEDTRUEFEATUREhelp.openai.com
  7. 07OpenAI's paid ChatGPT tiers get GPT-6 Sol.REPORTEDMOSTLY TRUEFEATUREopenai.com
  8. 08Claim removed during the check; its text is not republished.REPORTEDMIXEDATTRIBUTIONCUT FROM COPYmicrosoft.ai
  9. 09Anthropic's release page says Claude Haiku 5.5 costs "around 75% less to run" on average than its predecessor.VERIFIEDTRUEATTRIBUTIONanthropic.com
  10. 10Claim removed during the check; its text is not republished.FAILEDMOSTLY FALSEATTRIBUTIONCUT FROM COPYanthropic.com
  11. 11Anthropic's release page lists Haiku 5.5 at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens, and $0.50 and $2.50 above that. moelueker.com's launch review notes that, up to 100K tokens, that is exactly GPT-6 Luna's sticker price.REPORTEDMOSTLY TRUEPRICECORRECTED IN COPYanthropic.com
  12. 12Anthropic's release page lists Haiku 5.5 at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens, and $0.50 and $2.50 above that. moelueker.com's launch review notes that, up to 100K tokens, that is exactly GPT-6 Luna's sticker price.REPORTEDMOSTLY TRUEPRICECORRECTED IN COPYanthropic.com
  13. 13The second case is Anthropic's own price cut. ppc.land reported that the rate card fell from $1.00 and $5.00 on Haiku 4.5 to $0.10 and $0.50 on Haiku 5.5 for prompts up to 100,000 tokens, with larger prompts costing $0.50 and $2.50.REPORTEDMOSTLY TRUEPRICECORRECTED IN COPYanthropic.com
  14. 14Claude Haiku 5.5 has a long-prompt tier in which any request over 100,000 tokens costs five times more, at $0.50 per million input tokens and $2.50 per million output tokens.REPORTEDMOSTLY TRUEPRICEanthropic.com
  15. 15Claim removed during the check; its text is not republished.FAILEDFALSEBENCHMARKCUT FROM COPYvals.ai
  16. 16Claude Haiku 5.5 scored 46.4% on FrontierCode 1.1, according to Anthropic's comparison table.VERIFIEDTRUEBENCHMARKanthropic.com
  17. 17GPT-6 Luna scored 42.4% on FrontierCode 1.1, according to Anthropic's comparison table.REPORTEDMIXEDBENCHMARKdatalearner.com
  18. 18Aiidelist.com, reporting Anthropic's numbers, puts Claude Sonnet 5.5 at 70.6% on Terminal-Bench.REPORTEDMOSTLY TRUEBENCHMARKanthropic.com
  19. 19Artificial Analysis Intelligence Index version 4.1, released June 15, 2026, reweighted the index toward agentic tasks.VERIFIEDTRUEFEATUREartificialanalysis.ai
  20. 20Artificial Analysis Intelligence Index version 4.1, released June 15, 2026, added three new metrics: Cost per Task, Time per Task and Tokens per Task.VERIFIEDTRUEFEATUREartificialanalysis.ai
  21. 21Anthropic says halving Claude Sonnet 5.5's cache-read price makes Sonnet 5.5 around 20% cheaper on most agentic work.REPORTEDMOSTLY TRUEATTRIBUTIONanthropic.com
  22. 22Anthropic halved Claude Sonnet 5.5's cache-read price.VERIFIEDTRUEPRICEanthropic.com
  23. 23Anthropic cut Claude Sonnet 5.5's cache-read price on the same day it released Claude Haiku 5.5.VERIFIEDTRUEPRICEanthropic.com
  24. 24On June 15, 2026, Artificial Analysis announced version 4.1 of its Intelligence Index.REPORTEDMOSTLY TRUEHISTORYCORRECTED IN COPYartificialanalysis.ai
  25. 25Claude Haiku 5.5 beat GPT-6 Luna by 23.5 points on OSWorld 2.1 (72.4% minus 48.9%).COMPUTEDCOMPUTED
  26. 26Based on its 72.4% OSWorld score, Claude Haiku 5.5 needs about 1.38 attempts per success (1 divided by 0.724).COMPUTEDCOMPUTED
  27. 27Based on its 48.9% OSWorld score, GPT-6 Luna needs about 2.04 attempts per success (1 divided by 0.489).COMPUTEDCOMPUTED
  28. 28At an identical token price, GPT-6 Luna burns roughly 48% more attempts than Claude Haiku 5.5 for every finished task.COMPUTEDCOMPUTED
  29. 29Claude Haiku 5.5 leads GPT-6 Luna by 4 points on FrontierCode 1.1 (46.4% minus 42.4%).COMPUTEDCOMPUTED
  30. 30Claude Haiku 5.5's 39.2% on Terminal-Bench is a bit over half of Claude Sonnet 5.5's 70.6%.COMPUTEDCOMPUTED
  31. 31A job sending 1 million input tokens and receiving 200,000 output tokens on Claude Haiku 5.5, with requests under 100,000 tokens, costs $0.10 plus $0.10, or $0.20.COMPUTEDCOMPUTED
  32. 32The same 1 million input / 200,000 output token job on Claude Haiku 5.5 run through prompts over 100,000 tokens costs $0.50 plus $0.50, or $1.00.COMPUTEDCOMPUTED
  33. 33Assuming $0.20 per attempt, Claude Haiku 5.5 costs about $0.28 per finished task ($0.20 times 1.38).COMPUTEDCOMPUTED
  34. 34Assuming $0.20 per attempt, GPT-6 Luna costs about $0.41 per finished task ($0.20 times 2.04).COMPUTEDCOMPUTED
  35. 35Anthropic refreshed its entire model lineup (Opus 5.5, Sonnet 5.5, Haiku 5.5) in 15 days, from September 22 to October 7.COMPUTEDCOMPUTED
  36. 36If small-model running cost falls about 75% three more times by 2031, a $1.00 job becomes $0.25, then about $0.06, then about $0.016.COMPUTEDCOMPUTED
  37. 37Building a private test set of your own tasks costs a few dollars of tokens and a day of setup.COMPUTEDCOMPUTED

Every claim listed here was extracted from this article and checked against live sources before publication. The verdict is the checker's, not the writer's. Claims the check removed are counted but not republished.

Audit receipt KODA-20261009-36460207D9C6
Filed underToolsDeep Dive09 October 2026
Browse the Deep Dive archive

Get the morning Signal

196 editions so far, one a day. Unsubscribe anytime.