K Koda Intelligence
DEEP DIVE DEEP DIVE № 226 · 08 October 2026DOCKODA-20261008-2F86E4F6D1D0sha-256 of date + article + 24 checked + 13 computed

Same dime, two jobs:
price the prefix first

Anthropic now charges $0.10 04 per million tokens for fresh Claude Haiku 5.5 input, a tenth of Haiku 4.5's $1 13. On October 7 it cut cached Claude Sonnet 5.5 reads to that same dime, saying most agent work gets about 20% 02 cheaper. The parity covers input only. We think the first cost question is now which part of your prompt repeats on every call.

7 MIN READ · BY THE KODA EDITORIAL TEAM · PRICING · PROMPT CACHING
HAIKU 5.5 INPUT$0.10REPORTED CLAIM 04↓ 90% FROM $1
SONNET CACHE READ$0.10VERIFIED CLAIM 05↓ 50% FROM $0.20
SONNET 5.5 OUTPUT$10VERIFIED CLAIM 06Anthropic
HAIKU 5.5 INPUT$0.10↓ 90% FROM $1 SONNET CACHE READ$0.10↓ 50% FROM $0.20 SONNET 5.5 OUTPUT$10Anthropic HAIKU 5.5 OUTPUT$0.50Anthropic HAIKU CACHE READ$0.01Anthropic AGENT WORK SAVINGS20%ANTHROPIC CLAIM HAIKU 4.5 RETIRESOCT 15FORKAST

Anthropic now charges a dime for two very different jobs. Its pricing table lists fresh Claude Haiku 5.5 input at $0.10 04 per million tokens. On October 7, 2026 05, Anthropic also cut cached Claude Sonnet 5.5 reads from $0.20 to that same $0.10.

The small model reading new text now costs the same as the big model rereading old text. That moves the first cost question up a level. Before you ask which model is cheapest, ask which part of your prompt repeats on every call. The answer shapes your bill more than the model name does.

We will own the catch up front. The parity covers input only. Anthropic lists Sonnet 5.5 output at $10 25 per million tokens, which is 20 times Haiku 5.5's $0.50. Anthropic still says the cache cut makes Sonnet 5.5 about 20% 02 cheaper on most agent work.

The stakes reach well past one price sheet. GitHub made Haiku 5.5 generally available in Copilot on launch day. VentureBeat reported that Haiku 5.5's $0.10 03 input and $0.50 output match OpenAI's GPT-6 Luna only below 100,000 input tokens, as Haiku costs more above that threshold. Two big labs now sell small models at one price, so model price alone stops being an edge.

The Stable Prefix Decides the Bill

Every Claude request is billed on four line items. Fresh input, output, cache writes and cache reads each have their own rate on Anthropic's pricing table. A cache write stores a block of prompt text so later calls can reuse it. A cache read is that reuse, and it's the cheap line.

MONTHLY TOKEN BILL · OCTOBER 2026ANTHROPIC PRICING · KODA MATHBASE: 24 CHECKED + 13 COMPUTED, 4 SHOWN

One month of 100 million input and 10 million output tokens, priced four ways.

Fresh Sonnet 5.5 Anthropic list prices · before write fees COMPUTED CLAIM 30
$300
Cached Sonnet 5.5 Anthropic list prices · Sonnet output COMPUTED CLAIM 31
$110
Fresh Haiku 5.5 Anthropic list prices · under 100,000 tokens COMPUTED CLAIM 32
$15
Cached Haiku 5.5 Anthropic list prices · $0.01 reads COMPUTED CLAIM 32
$6

The pattern that wins on this table is a stable-prefix, variable-suffix design. The prefix is the opening block of the prompt that never changes: system instructions, tool definitions, policy manuals, product catalogs, a repository snapshot. The suffix is everything that does change, such as the user's question or the tool result from the last turn. You cache the prefix once and pay fresh rates only on the suffix.

Anthropic's documentation says a cache hit on Sonnet 5.5 costs 5% 19 of the standard $2.00 09 input price. Anthropic prices a five-minute write at $2.50 per million, a 1.25 multiplier on that input price. On Anthropic's numbers, the write overhead is $0.50 26 per million and each hit saves $1.90. A single reuse covers the write.

From there, the table sorts every model choice by where its tokens sit. Cached Sonnet input at $0.10 04 matches fresh Haiku input at $0.10, per Anthropic's list. Haiku cache reads, which Anthropic lists at $0.01 10 per million tokens for prompts up to 100,000 tokens, are ten times cheaper than either. So decide the prefix first. The model decision follows, piece by piece, for the suffix.

Where the Margin Actually Leaks

Why would a company halve the price of something it already sells? Unite.AI reports that Anthropic said cache reads make up a large share of token consumption. The discount lands where the volume already is. That tells you where Anthropic thinks your money goes.

Take a 100,000-token 27 context that returns 2,000 output tokens, priced at Anthropic's list rates. By our math, fresh Haiku costs $0.01 for input and $0.001 for output, so $0.011 a call. Cached Sonnet costs $0.01 28 for input and $0.02 for output, so $0.03. The input lines are identical, and Sonnet still costs about 2.7 times as much.

Stretch the answers to 20,000 29 output tokens and fresh Haiku costs $0.02 while cached Sonnet costs $0.21. That's a 10.5 times gap. Output is the line that eats the margin.

Now scale it to a month of 100 million 30 input tokens and 10 million output tokens on Anthropic's list prices. Fresh Sonnet runs $300 before any write fees. Cached Sonnet input with Sonnet output runs $110 31. Fresh Haiku runs $15 32, and cached Haiku runs $6.

The input headline is a distraction. Founders see $0.10 20 next to $0.10 and start arguing about models. But the asset that produces the output is the prefix you wrote once. Build it well and every model gets cheaper to run on top of it.

The parity also breaks above 100,000 tokens 11. Anthropic lists Haiku 5.5 at $0.50 input and $2.50 output past that line. A fresh one-million-token Haiku call with 10,000 33 output tokens then costs about $0.525, while cached Sonnet costs about $0.20. Cache the Haiku prefix too, at Anthropic's $0.05 34 long-context read rate, and that call drops to about $0.075.

One damaging admission on that Sonnet figure: Anthropic's launch table shows no separate long-context Sonnet rate. Treat the $0.20 05 as a floor. Quality is the bigger gap. A cheap answer that fails is the most expensive answer you can sell.

Anthropic reports Sonnet 5.5 at 70.6% 15 on Terminal-Bench 4.0. Forkast reports Haiku 5.5 at 72.4% 16 on OSWorld, up from 15.7% 17 for its predecessor. Those are different tests, so neither number tells you which model finishes your task. Alex Wang, who works in applied AI at Rogo, gave Anthropic a statement that VentureBeat printed.

"It's accurate enough that we'd trust it there and fast and cheap enough that we can run it a lot."

If every Haiku answer needs a person to check it, you're running a service business dressed up as software. Those human minutes cost more than any token line. Count cost per finished task, and the right model often surprises you. Lower cost per task widens gross margin, and that raises what you can afford to pay to win a customer against their lifetime value.

It is unclear whether your traffic looks like the "most agentic work" behind Anthropic's 20% 02 claim. That figure is Anthropic's own, measured on workloads Anthropic chose. Nobody outside your logs knows your hit rate. One flag sits outside pricing entirely: whether tenant or permission data belongs in a cache is a security call.

The work that fixes the bill is plain. Log tokens, move the stable text to the top, read the bill, repeat. Nobody posts screenshots of it. It's still where the margin comes from.

What the matching dime leaves off the bill

OUTPUT GAP
20×

Sonnet output still costs 20 25 times Haiku's.

Anthropic lists Sonnet 5.5 output at $10 06 per million tokens against Haiku 5.5's $0.50 07. On a 100,000-token 29 call with 20,000 output tokens, fresh Haiku costs $0.02 and cached Sonnet costs $0.21, a 10.5 times gap.

LONG CONTEXT
$0.525 33

Parity breaks above 100,000 tokens 11.

Past that line Anthropic lists Haiku 5.5 at $0.50 input and $2.50 output, so a fresh one-million-token call with 10,000 output tokens costs about $0.525. Caching the Haiku prefix at the $0.05 34 long-context read rate drops it to about $0.075.

CONTEXT ASSET
2031

The cached prefix keeps its lead as prices fall.

Our slow assumption halves the small-model input floor each year to about $0.003 36 by 2031. The token line shrinks toward noise, and the prefix plus the logs showing which context gets reused become the scarce asset.

2031: Context Becomes the Billable Asset

The dated prices trace a steep line. Forkast puts Haiku 4.5 at $1 13 input and $5 output, and Anthropic lists Haiku 5.5 at a tenth of both. Anthropic said on September 22 that Opus 5.5 costs 40% 24 less to run than Opus 5. Sonnet 5.5 launched September 28, and nine days later Anthropic halved its cache reads.

Forkast reports that Haiku 4.5 retires October 15, eight days after its replacement shipped. Forkast also puts the average open-weights price on OpenRouter at $0.83 per million tokens, with Mistral ML4 at $0.68 22 input. Anthropic's new floor is already below most of the market it competes with.

Our projection is an assumption, and a slow one. Say the small-model input floor halves each year from today's $0.10 20. That's far slower than the 90% cut VentureBeat reported for Haiku 5.5 prompts up to 100,000 tokens (50% above that). It gives $0.05 36 in 2027, about $0.0125 in 2029 and about $0.003 in 2031. At that rate, 100 million 37 input tokens a month costs about 31 cents.

At that price, the token line stops being a business decision. What stays scarce is the prefix itself, along with the logs that show which context gets reused. Those compound, since each week of logs makes the next prefix tighter. Matching GPT-6 Luna's price gives Anthropic no lasting edge, so we expect labs to compete on cache rules next.

The risk is lopsided. Restructuring a prompt into a stable prefix costs some engineering time, once. Anthropic prices cached reads at a small fraction of fresh input on both models, so the cached design keeps its lead as prices fall. We think builders who treat context as an asset now will hold the cheapest cost base in 2031, whichever lab wins the price war.

Reprice One Agent From Real Logs

Pick one agent with steady traffic, like a support bot or a code reviewer. The job is to price it on four line items and then route it.

Step one is exporting seven days of requests. For each call, record input tokens and output tokens in a spreadsheet. Note which calls ran on Haiku 5.5 and which ran on Sonnet 5.5.

Step two is finding the prefix, the opening text that stays identical across calls. Move every changing piece below it, including timestamps and user names. A timestamp at the top quietly breaks every cache match.

Step three is the four-line math. Anthropic's list prices for Sonnet 5.5 are $2.00 09 fresh input, $10 06 output, $2.50 five-minute write and $0.10 04 read. For Haiku 5.5 under 100,000 tokens 14, Anthropic lists $0.10, $0.50 11, $0.125 and $0.01 10. Multiply each token count by its rate, add the four lines and compare both models.

Step four is turning caching on for the prefix for one day. Divide the cached-read tokens on your usage report by the prefix tokens you sent. That ratio is your hit rate, the share of prefix tokens served from storage. Expect the first number to disappoint, because a reordered tool list or a five-minute gap between calls kills hits.

Step five is routing. Send every call to Haiku 5.5 first, and escalate to Sonnet 5.5 only when Haiku fails a check you wrote, like invalid JSON. Cap Sonnet's answers with a schema and a length limit, since Anthropic bills its output at $10 06 per million.

Rerun the spreadsheet next Friday with fresh logs. If cost per finished task did not drop, the prefix is still leaking, so go back to step two. Something will break on the first pass, and the logs show you exactly where.

DOJO · BUILD THIS WEEKEND

Reprice one steady agent from its real logs.

  1. Export seven days of requests. Record input and output tokens per call in a spreadsheet and tag each call as Haiku 5.5 or Sonnet 5.5.
  2. Move the stable prefix to the top. Put system instructions and tool definitions first, and push timestamps and user names below them, since a timestamp at the top breaks every cache match.
  3. Route Haiku first and cap Sonnet. Escalate to Sonnet 5.5 only when Haiku fails a check you wrote, like invalid JSON, and limit Sonnet answers with a schema because its output bills at $10 25 per million tokens.
Practice: Design an AI-Assisted Workflow
THE BOTTOM LINE

Write the prefix once and every model gets cheaper to run.

Fresh Haiku and cached Sonnet now share a $0.10 04 input price, so model price alone stops being an edge. Output, long context and human review still decide the real bill. Count cost per finished task, measure your own hit rate, and treat Anthropic's 20% 02 claim as a claim. We expect the builders who log, reorder and reread their bills now to hold the cheapest cost base whichever lab wins the price war.

LISTEN · AUDIO BRIEFINGThe conversation · ~17 min
WATCH · VISUAL NARRATIVEAnimated breakdown · ~6 min
PLAY · YOUTUBE
EDITORIAL RECEIPTKODA-20261008-2F86E4F6D1D0
As of08 October 2026MethodClaim extraction, dated-evidence review, and temporal consistency gate.CorrectionsContact the Koda desk
EVIDENCE24 CHECKED + 13 COMPUTED · 12 VERIFIED · 12 REPORTED
12 verified12 reported13 computed
  1. 01GitHub made Claude Haiku 5.5 generally available in GitHub Copilot on Haiku 5.5's launch day.VERIFIEDTRUEFEATUREgithub.blog
  2. 02Anthropic says the cache-read price cut makes Claude Sonnet 5.5 about 20% cheaper on most agent work.VERIFIEDTRUEATTRIBUTIONanthropic.com
  3. 03VentureBeat reported that Haiku 5.5's $0.10 input and $0.50 output match OpenAI's GPT-6 Luna only below 100,000 input tokens, as Haiku costs more above that threshold.REPORTEDMOSTLY TRUEATTRIBUTIONCORRECTED IN COPYventurebeat.com
  4. 04Anthropic's pricing table lists fresh (uncached) Claude Haiku 5.5 input at $0.10 per million tokens.REPORTEDMOSTLY TRUEPRICEanthropic.com
  5. 05On October 7, 2026, Anthropic cut the price of cached Claude Sonnet 5.5 reads from $0.20 to $0.10 per million tokens.VERIFIEDTRUEPRICEplatform.claude.com
  6. 06Anthropic lists Claude Sonnet 5.5 output at $10 per million tokens.VERIFIEDTRUEPRICEsupport.claude.com
  7. 07Anthropic lists Claude Haiku 5.5 output at $0.50 per million tokens.REPORTEDMOSTLY TRUEPRICEanthropic.com
  8. 08Claude Sonnet 5.5's standard (fresh) input price is $2.00 per million tokens.VERIFIEDTRUEPRICEanthropic.com
  9. 09Anthropic prices a five-minute cache write for Claude Sonnet 5.5 at $2.50 per million tokens, a 1.25 multiplier on the $2.00 input price.VERIFIEDTRUEPRICEanthropic.com
  10. 10Haiku cache reads, which Anthropic lists at $0.01 per million tokens for prompts up to 100,000 tokens, are ten times cheaper than either.REPORTEDMOSTLY TRUEPRICECORRECTED IN COPYanthropic.com
  11. 11Anthropic lists Claude Haiku 5.5 at $0.50 input and $2.50 output per million tokens for prompts above 100,000 tokens.VERIFIEDTRUEPRICEanthropic.com
  12. 12Anthropic's long-context cache read rate for Claude Haiku 5.5 is $0.05 per million tokens.VERIFIEDTRUEPRICEanthropic.com
  13. 13Forkast puts Claude Haiku 4.5 pricing at $1 input and $5 output per million tokens.VERIFIEDTRUEPRICEdocs.anthropic.com
  14. 14Anthropic lists a five-minute cache write for Claude Haiku 5.5 (under 100,000 tokens) at $0.125 per million tokens.VERIFIEDTRUEPRICEanthropic.com
  15. 15Anthropic reports Claude Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0.REPORTEDMOSTLY TRUEBENCHMARKanthropic.com
  16. 16Forkast reports Claude Haiku 5.5 scores 72.4% on OSWorld.REPORTEDMOSTLY TRUEBENCHMARKkingy.ai
  17. 17Forkast reports Claude Haiku 5.5's predecessor scored 15.7% on OSWorld.REPORTEDMOSTLY TRUEBENCHMARKtestingcatalog.com
  18. 18Forkast reports that Claude Haiku 4.5 retires on October 15, 2026, eight days after its replacement Claude Haiku 5.5 shipped.REPORTEDMIXEDMODELsupport.claude.com
  19. 19Anthropic's documentation says a cache hit on Claude Sonnet 5.5 costs 5% of the standard input price.VERIFIEDTRUEATTRIBUTIONdocs.anthropic.com
  20. 20Say the small-model input floor halves each year from today's $0.10, far slower than the 90% cut VentureBeat reported for Haiku 5.5 prompts up to 100,000 tokens (50% above that).REPORTEDMOSTLY TRUEATTRIBUTIONCORRECTED IN COPYventurebeat.com
  21. 21Anthropic's launch pricing table for Claude Sonnet 5.5 shows no separate long-context rate.REPORTEDMOSTLY TRUEPRICEsupport.claude.com
  22. 22Forkast reports Mistral ML4 input is priced at $0.68 per million tokens.REPORTEDMIXEDPRICEfourweekmba.com
  23. 23Claude Sonnet 5.5 launched on September 28, 2026, and Anthropic halved its cache read price nine days later.VERIFIEDTRUEHISTORYanthropic.com
  24. 24Anthropic said on September 22, 2026 that Claude Opus 5.5 costs 40% less to run than Claude Opus 5.REPORTEDMOSTLY TRUESTATanthropic.com
  25. 25Claude Sonnet 5.5's output price of $10 per million tokens is 20 times Claude Haiku 5.5's output price of $0.50 per million tokens.COMPUTEDCOMPUTED
  26. 26On Anthropic's Claude Sonnet 5.5 prices, the cache write overhead is $0.50 per million tokens and each cache hit saves $1.90 per million tokens, so one reuse pays for the write.COMPUTEDCOMPUTED
  27. 27For a 100,000-token context returning 2,000 output tokens at list rates, fresh Claude Haiku 5.5 costs $0.01 for input and $0.001 for output, totaling $0.011 per call.COMPUTEDCOMPUTED
  28. 28For a 100,000-token context returning 2,000 output tokens at list rates, cached Claude Sonnet 5.5 costs $0.01 for input and $0.02 for output, totaling $0.03 per call, about 2.7 times fresh Claude Haiku 5.5.COMPUTEDCOMPUTED
  29. 29For a 100,000-token context returning 20,000 output tokens, fresh Claude Haiku 5.5 costs $0.02 per call while cached Claude Sonnet 5.5 costs $0.21, a 10.5-times gap.COMPUTEDCOMPUTED
  30. 30For a month of 100 million input tokens and 10 million output tokens at Anthropic list prices, fresh Claude Sonnet 5.5 costs $300 before write fees.COMPUTEDCOMPUTED
  31. 31For a month of 100 million input tokens and 10 million output tokens, cached Claude Sonnet 5.5 input with Sonnet output costs $110.COMPUTEDCOMPUTED
  32. 32For a month of 100 million input tokens and 10 million output tokens, fresh Claude Haiku 5.5 costs $15 and cached Claude Haiku 5.5 costs $6.COMPUTEDCOMPUTED
  33. 33A fresh one-million-token Claude Haiku 5.5 call with 10,000 output tokens costs about $0.525, while the same call on cached Claude Sonnet 5.5 costs about $0.20.COMPUTEDCOMPUTED
  34. 34A one-million-token Claude Haiku 5.5 call with 10,000 output tokens and a cached prefix at the $0.05 long-context read rate costs about $0.075.COMPUTEDCOMPUTED
  35. 35Claude Haiku 5.5's input and output prices are one tenth of Claude Haiku 4.5's prices.COMPUTEDCOMPUTED
  36. 36If the small-model input price floor halves each year from $0.10 per million tokens, it reaches $0.05 in 2027, about $0.0125 in 2029 and about $0.003 in 2031.COMPUTEDCOMPUTED
  37. 37At a projected 2031 price of about $0.003 per million tokens, 100 million input tokens a month costs about 31 cents.COMPUTEDCOMPUTED

Every claim listed here was extracted from this article and checked against live sources before publication. The verdict is the checker's, not the writer's. Claims the check removed are counted but not republished.

Audit receipt KODA-20261008-2F86E4F6D1D0
Filed underPricingDeep Dive08 October 2026
Browse the Deep Dive archive

Get the morning Signal

195 editions so far, one a day. Unsubscribe anytime.