K Koda Intelligence
DEEP DIVE DEEP DIVE № · 19 September 2026DOCKODA-20260919-CCB644249BA7sha-256 of date + article + 24 claims
FILED 19 SEPTEMBER 202624 CLAIMS CHECKED · 11 VERIFIED

Claude leads 26% of its own R&D.
The other 64 points are the story.

Anthropic's September 17 report says Claude now leads about 26%12 of the company's model research work, up from under 1% in February 2026. The number the headlines skipped: Claude participates in more than 90% of the same tasks at a collaborate level or higher. That 64-point spread is where the real work sits. With Claude Fable 5.101 and GPT-6 Astra tied at 53 on the Intelligence Index v4.3, gains now come from workflow, not raw IQ.

6 MIN READ · BY THE KODA EDITORIAL TEAM · STRATEGY · AGENT WORKFLOWS
AI-LED R&D WORK26%REPORTED CLAIM 12↑ 25PP SINCE FEB
CLAUDE PARTICIPATES90%NOT MEASUREDANTHROPIC
OWNERSHIP GAP64 PTSNOT MEASUREDAL3 TO AL4
AI-LED R&D WORK26%↑ 25PP SINCE FEB CLAUDE PARTICIPATES90%ANTHROPIC OWNERSHIP GAP64 PTSAL3 TO AL4 CONCURRENT AGENTS30,000INTERNAL PLATFORM BLOCKED DECISIONS1 IN 47,000AUGUST SAFETY COMPUTE, AI-LED12%↑ 6PP VS BASELINE CODE SHIPPED8XVS 2021 TO 2025 AI RELEASES LOGGED50THURSDAI SEPT 2026

Anthropic says Claude now leads 26%12 of the research work that builds the next Claude. In February 2026 that number was reportedly under 1%. The report landed on September 17, 2026, from the Anthropic Institute. Every headline ran the 26%.

Here is the number the headlines skipped. Claude touches more than 90% of the same work at a "collaborates" level or higher. So the model helps with almost everything and owns about a quarter. That 64-point spread is the whole story.

I will own what I do not know up front. Every figure here is self-measured by Anthropic. A Claude model read Slack records and internal docs and graded Claude's own work.

So this is not a piece about robots building robots. It is a piece about how one lab rebuilt knowledge work so an agent could hold a ticket end to end. And it is a warning about what happens when you turn that into a number on a dashboard.

The Ownership Gap

Name the thing. Anthropic scored its work on an Automation Level scale built by Epoch AI, an independent nonprofit. AL0 means no AI. AL5 means the AI runs the whole thing with no human in the loop.

OWNERSHIP LEDGER · SEPTEMBER 2026ANTHROPIC INSTITUTE · EPOCH AI LADDER · THURSDAIBASE: 24 CHECKED CLAIMS, 4 SHOWN

What it actually costs to move a task from assist to ownership.

AI-led task share Anthropic Institute · AL4 or higher REPORTED CLAIM 12
26%
Assist-level coverage Anthropic Institute · AL3 or higher NOT MEASURED
90%
Safety compute inside AI-led work Anthropic · doubles from 6% baseline NOT MEASURED
12%
Transcripts reaching a human Monitor funnel · per week NOT MEASURED
50

Two rungs matter. AL3 is "collaborates": the AI does big chunks of a task while a human directs closely. AL4 is "leads": the AI takes a high-level prompt and carries most of the task to completion while a human supervises. Anthropic reports zero work at AL5.

I call the distance between those two rungs the Ownership Gap. Assist is cheap. Ownership is expensive. Getting a model to help on 90% of tasks takes a subscription and a Slack channel. Getting it to own 26%12 takes a rebuilt workflow.

Think of it as three buckets, not a spectrum. Bucket one is work a human still drives, under 10% at Anthropic. Bucket two is work the agent assists on but cannot finish alone, at least 64 points and by far the biggest bucket. Bucket three is the 26%12 the agent finishes from a prompt.

Progress is moving work from bucket two to bucket three. That is a plumbing job, not a model job. The benchmark data backs this up: Claude Fable 5.101 and GPT-6 Astra sit tied at 53 on the Artificial Analysis Intelligence Index v4.3. When raw model gains flatten at the top, the gains have to come from how you structure the work.

Tickets Claude Can Actually Close

Here is the deep dive, builder to builder. An AI agent is a 500 IQ intern. Brilliant, tireless, zero context on day one. An intern cannot lead a task you have not written down, and Anthropic's 26%12 came from writing the tasks down.

Amateurs count how much the agent does; leaders count where the agent does it. A hundred agent-led log-parsing tickets carry less risk than one agent-led evaluation design.· KODA EDITORIAL · SEPTEMBER 2026

Look at how they built the catalogue. Anthropic pulled work records from Slack and internal documentation and assembled a bottom-up list of R&D tasks. Then, for each week of July 2026, a Claude judge read the evidence and rated how automated each task category was. You cannot score what you have not itemized. Step one was inventory.

That is the 80/20 of this whole report. Twenty percent of the effort is model quality. Eighty percent is turning fuzzy research into tickets with a clear prompt, a clear output, and a clear check. "Design, run, and analyze a sweep of experiments" is a ticket an agent can lead. "figure out why the loss curve looks weird" is not, yet.

Now the plumbing underneath it. Anthropic says roughly 30,000 agents run at once on its most-used internal platform, and every action there passes through an online monitor. Of more than 1 billion agent decisions in August, about 1 in 47,000 got blocked. Roughly 100,000 transcripts a week get flagged and about 50 a week reach a human.

That funnel is the real product. An ounce in pre is worth a pound in post: the monitor, the logging, and the ticket templates existed before the 26%12 could. Simple beats complex here, every time. The tasks that moved to "leads" were the ones with a boring, repeatable shape.

Translate it to dollars. Anthropic says its engineers now ship 8x the code per quarter they did across 2021 to 2025. Run your own napkin math: if a researcher costs $300,000 a year and an agent owns 26%12 of their tickets, that is $78,000 of execution moved to an agent. The model did not get 26% smarter. The work got 26% more legible.

But nobody should read this as free. Anthropic says 6% of its AI R&D compute goes to safety work, and inside the AI-led slice that share doubles to 12%. Ownership costs more oversight per task, not less. Anthropic's own numbers say so.

What the 64-point spread actually hides

OWNERSHIP MOVED
26%12

The jump came from plumbing, not IQ.

Anthropic built a bottom-up task catalogue from Slack and internal docs, then rated each category week by week for July 2026. With Claude Fable 5.101 and GPT-6 Astra tied at 53 on the Intelligence Index v4.3, the leverage is in ticket shape, not model gains.

OVERSIGHT COST
12%

Ownership buys more supervision, not less.

Anthropic says 6% of its AI R&D compute goes to safety work, and inside the AI-led slice that share doubles to 12%. Roughly 30,000 agents run at once behind an online monitor that flags about 100,000 transcripts a week.

METRIC DRIFT
2031

The score is a Claude judge grading Claude.

Every figure is self-measured, and any team can shard tasks smaller or relabel collaborates as leads while nothing real changes. If share of AI-led work becomes a standard line item by 2031, Goodhart's Law arrives with it.

2031: Every Team Publishes a Ladder Score

Pull back five years. Git went from a hacker tool to a hiring requirement in about a decade, and CI pipelines followed the same path. My read is that "share of AI-led work" rides the same arc. By 2031 it is a line item most knowledge teams report, whether or not the number means much.

That is the asymmetric bet. If you instrument your workflow now and the metric turns out to be noise, you lost a spreadsheet. If you skip it and the metric becomes the yardstick, you are five years behind on task hygiene. Headcount buys output. A task catalogue buys leverage.

Here is the warning half. Goodhart's Law says a measure that becomes a target stops being a good measure. Anthropic's 26%12 is a Claude judge grading Claude. Any team can shard tasks smaller or relabel "collaborates" as "leads," and the dashboard climbs while nothing real changes.

I cannot tell whether the 26%12 measures autonomy or just execution throughput. Anthropic says humans there still set goals, gate safety, and decide what ships. Claude does not pick research directions or reallocate compute. The metric tells you where the intern sits, not who runs the lab.

Two contrasts worth holding onto. Amateurs count how much the agent does; leaders count where the agent does it. A hundred agent-led log-parsing tickets carry less risk than one agent-led evaluation design. I think the Ownership Gap needs a second axis, stakes, before it means anything outside Anthropic's walls.

Impermanence cuts both ways here. Anthropic frames the index as a prototype it expects to revise. Dario Amodei's September 12 essay asked the industry to accept embedded third-party evaluators, and this post reads like the data those evaluators would check. An Anthropic worker publicly quit on September 9 over risk concerns, so the same lab produced both stories in the same month.

Score Your Own Task Catalogue by Friday

Show, don't tell. You do not need a frontier lab or a CS degree to run this exercise. You need a spreadsheet, one week of Slack history, and an honest hour.

First, inventory. Pull every recurring task your team did last week and aim for 30 to 50 line items. "Wrote the weekly metrics summary" counts. "Thought about strategy" does not.

Second, rate each one on the Epoch AI ladder. AL0 means no AI touched it. AL3 means an agent did big chunks under your direction, and AL4 means you gave a prompt and only reviewed the output. Be brutal, because the whole value is in the honesty.

Third, compute your own Ownership Gap. Take the percent at AL3 or higher and subtract the percent at AL4. Expect a wide gap. Assist is everywhere and ownership is rare, and that is normal on a first pass.

Fourth, pick three bucket-two tasks and rewrite them as tickets with a clear input, output, and check. Hand one to an agent and let it break, because it will break the first time. Log why it broke. That log is your monitor funnel in miniature.

Two tool notes from today's digest. ThursdAI's September 2026 ledger counts 50 AI releases through September 17, 11 of them last week. Read the ledger, not the launch posts, and only add a "new model" row to your sheet when a rating actually moves.

Same skepticism applies to tool directories. The Featured slot on the Rosebud AI listing is reportedly a paid Hola AI placement carrying a ref=featured tag, not a ranking. A dashboard number is only as good as who scored it, and that applies to Anthropic's 26%12 too.

Get your reps in. Rerun the sheet in four weeks. If your AL4 share moved, you learned what ticket shape your agents can own. If it did not move, you learned that too, and learning it is the whole point of measuring.

DOJO · BUILD THIS WEEKEND

Score your own task catalogue before Friday.

  1. Inventory one week of real work. Pull every recurring task your team did last week from Slack and docs, and aim for 30 to 50 line items. "Wrote the weekly metrics summary" counts; "thought about strategy" does not.
  2. Rate each line on the Epoch AI ladder. AL0 means no AI touched it, AL3 means an agent did big chunks under your direction, AL4 means you gave a prompt and only reviewed output. Then subtract your AL4 share from your AL3-or-higher share to get your own Ownership Gap.
  3. Rewrite three assist-only tasks as tickets. Give each a clear input, output and check, hand one to an agent, and log exactly why it breaks the first time. Rerun the whole sheet in four weeks and see whether your AL4 share actually moved.
Train the full skill in The Dojo
THE BOTTOM LINE

The 26%12 is a plumbing score, not an autonomy score.

Anthropic did not discover a model that builds models; it rebuilt its own research work until an agent could hold a ticket end to end. The 64-point gap between 90% participation and 26% ownership is the part any team can copy, because it is made of inventories, prompts, checks and monitors rather than frontier weights. The costs travel with it: safety compute doubles to 12% inside the AI-led slice, and about 50 flagged transcripts a week still land on a human. Treat the number as a diagnostic of task legibility, not a leaderboard, and remember it was graded by a Claude judge reading Claude's work. Score your own catalogue, then argue with the result in four weeks.

LISTEN · AUDIO BRIEFINGThe conversation · ~15 min
WATCH · VISUAL NARRATIVEAnimated breakdown · ~9 min
PLAY · YOUTUBE
EDITORIAL RECEIPTKODA-20260919-CCB644249BA7
As of19 September 2026MethodClaim extraction, dated-evidence review, and temporal consistency gate.CorrectionsContact the Koda desk
EVIDENCE24 CLAIMS CHECKED · 11 VERIFIED · 13 REPORTED
11 verified13 reported
  1. 01Claude Fable 5.1 and GPT-6 Astra are tied at a score of 53 on the Artificial Analysis Intelligence Index v4.3.VERIFIEDTRUEBENCHMARK
  2. 02Claude Fable 5.1 is an AI model listed on the Artificial Analysis Intelligence Index v4.3.VERIFIEDTRUEMODEL
  3. 03GPT-6 Astra is an AI model listed on the Artificial Analysis Intelligence Index v4.3.VERIFIEDTRUEMODEL
  4. 04On the Epoch AI Automation Level scale, AL0 means no AI was involved.VERIFIEDTRUEFEATURE
  5. 05On the Epoch AI Automation Level scale, AL5 means the AI runs the whole task with no human in the loop.VERIFIEDTRUEFEATURE
  6. 06On the Epoch AI Automation Level scale, AL3 is labeled "collaborates" and means the AI does big chunks of a task while a human directs closely.VERIFIEDTRUEFEATURE
  7. 07On the Epoch AI Automation Level scale, AL4 is labeled "leads" and means the AI takes a high-level prompt and carries most of the task to completion while a human supervises.VERIFIEDTRUEFEATURE
  8. 08Every agent action on Anthropic's internal agent platform passes through an online monitor.REPORTEDMOSTLY TRUEFEATURECORRECTED IN COPY
  9. 09Humans at Anthropic still set research goals, gate safety, and decide what ships, rather than Claude.REPORTEDMOSTLY TRUEFEATURECORRECTED IN COPY
  10. 10Claude does not pick research directions or reallocate compute at Anthropic.VERIFIEDTRUEFEATURE
  11. 11The Featured slot on the Rosebud AI tool directory listing is a paid Hola AI placement carrying a ref=featured tag, not a ranking.REPORTEDUNVERIFIABLEFEATURECORRECTED IN COPY
  12. 12The Anthropic Institute published a report on September 17, 2026 stating that Claude leads 26% of the research work building the next Claude.REPORTEDMOSTLY TRUEATTRIBUTION
  13. 13Anthropic's automation figures were self-measured, with a Claude model reading Slack records and internal documents and grading Claude's own work.REPORTEDMOSTLY TRUEATTRIBUTION
  14. 14Claim removed during the check; its text is not republished.REPORTEDMIXEDATTRIBUTIONCUT FROM COPY
  15. 15Anthropic scored its research work using an Automation Level scale built by Epoch AI, an independent nonprofit.REPORTEDMOSTLY TRUEATTRIBUTION
  16. 16Anthropic pulled work records from Slack and internal documentation to assemble a bottom-up list of R&D tasks for its automation report.VERIFIEDTRUEATTRIBUTION
  17. 17For each week of July 2026, a Claude judge rated how automated each Anthropic R&D task was, based on Slack and internal documentation evidence.REPORTEDMOSTLY TRUEATTRIBUTIONCORRECTED IN COPY
  18. 18Goodhart's Law states that a measure that becomes a target stops being a good measure.VERIFIEDTRUEATTRIBUTION
  19. 19Anthropic frames its Automation Level index of Claude-led research work as a prototype it expects to revise.REPORTEDMOSTLY TRUEATTRIBUTION
  20. 20Dario Amodei published an essay on September 12, 2026 asking the AI industry to accept embedded third-party evaluators.VERIFIEDTRUEATTRIBUTION
  21. 21Git went from a hacker tool to a hiring requirement in about a decade.REPORTEDMOSTLY TRUEHISTORY
  22. 22An Anthropic worker publicly quit on September 9, 2026 over risk concerns.REPORTEDMOSTLY TRUEHISTORY
  23. 23Anthropic says Claude now leads 26% of the research work that builds the next Claude model.REPORTEDMOSTLY TRUESTAT
  24. 24In February 2026, the share of Anthropic's research work led by Claude was under 1%.REPORTEDMOSTLY TRUESTATCORRECTED IN COPY

Every claim listed here was extracted from this article and checked against live sources before publication. The verdict is the checker's, not the writer's. Claims the check removed are counted but not republished.

Audit receipt KODA-20260919-CCB644249BA7
Filed underStrategyDeep Dive19 September 2026
Browse the Deep Dive archive

Get the morning Signal

176 editions so far, one a day. Unsubscribe anytime.