Anthropic says Claude now leads 26%12 of the research work that builds the next Claude. In February 2026 that number was reportedly under 1%. The report landed on September 17, 2026, from the Anthropic Institute. Every headline ran the 26%.
Here is the number the headlines skipped. Claude touches more than 90% of the same work at a "collaborates" level or higher. So the model helps with almost everything and owns about a quarter. That 64-point spread is the whole story.
I will own what I do not know up front. Every figure here is self-measured by Anthropic. A Claude model read Slack records and internal docs and graded Claude's own work.
So this is not a piece about robots building robots. It is a piece about how one lab rebuilt knowledge work so an agent could hold a ticket end to end. And it is a warning about what happens when you turn that into a number on a dashboard.
The Ownership Gap
Name the thing. Anthropic scored its work on an Automation Level scale built by Epoch AI, an independent nonprofit. AL0 means no AI. AL5 means the AI runs the whole thing with no human in the loop.
What it actually costs to move a task from assist to ownership.
Two rungs matter. AL3 is "collaborates": the AI does big chunks of a task while a human directs closely. AL4 is "leads": the AI takes a high-level prompt and carries most of the task to completion while a human supervises. Anthropic reports zero work at AL5.
I call the distance between those two rungs the Ownership Gap. Assist is cheap. Ownership is expensive. Getting a model to help on 90% of tasks takes a subscription and a Slack channel. Getting it to own 26%12 takes a rebuilt workflow.
Think of it as three buckets, not a spectrum. Bucket one is work a human still drives, under 10% at Anthropic. Bucket two is work the agent assists on but cannot finish alone, at least 64 points and by far the biggest bucket. Bucket three is the 26%12 the agent finishes from a prompt.
Progress is moving work from bucket two to bucket three. That is a plumbing job, not a model job. The benchmark data backs this up: Claude Fable 5.101 and GPT-6 Astra sit tied at 53 on the Artificial Analysis Intelligence Index v4.3. When raw model gains flatten at the top, the gains have to come from how you structure the work.
Tickets Claude Can Actually Close
Here is the deep dive, builder to builder. An AI agent is a 500 IQ intern. Brilliant, tireless, zero context on day one. An intern cannot lead a task you have not written down, and Anthropic's 26%12 came from writing the tasks down.
Look at how they built the catalogue. Anthropic pulled work records from Slack and internal documentation and assembled a bottom-up list of R&D tasks. Then, for each week of July 2026, a Claude judge read the evidence and rated how automated each task category was. You cannot score what you have not itemized. Step one was inventory.
That is the 80/20 of this whole report. Twenty percent of the effort is model quality. Eighty percent is turning fuzzy research into tickets with a clear prompt, a clear output, and a clear check. "Design, run, and analyze a sweep of experiments" is a ticket an agent can lead. "figure out why the loss curve looks weird" is not, yet.
Now the plumbing underneath it. Anthropic says roughly 30,000 agents run at once on its most-used internal platform, and every action there passes through an online monitor. Of more than 1 billion agent decisions in August, about 1 in 47,000 got blocked. Roughly 100,000 transcripts a week get flagged and about 50 a week reach a human.
That funnel is the real product. An ounce in pre is worth a pound in post: the monitor, the logging, and the ticket templates existed before the 26%12 could. Simple beats complex here, every time. The tasks that moved to "leads" were the ones with a boring, repeatable shape.
Translate it to dollars. Anthropic says its engineers now ship 8x the code per quarter they did across 2021 to 2025. Run your own napkin math: if a researcher costs $300,000 a year and an agent owns 26%12 of their tickets, that is $78,000 of execution moved to an agent. The model did not get 26% smarter. The work got 26% more legible.
But nobody should read this as free. Anthropic says 6% of its AI R&D compute goes to safety work, and inside the AI-led slice that share doubles to 12%. Ownership costs more oversight per task, not less. Anthropic's own numbers say so.
What the 64-point spread actually hides
The jump came from plumbing, not IQ.
Anthropic built a bottom-up task catalogue from Slack and internal docs, then rated each category week by week for July 2026. With Claude Fable 5.101 and GPT-6 Astra tied at 53 on the Intelligence Index v4.3, the leverage is in ticket shape, not model gains.
Ownership buys more supervision, not less.
Anthropic says 6% of its AI R&D compute goes to safety work, and inside the AI-led slice that share doubles to 12%. Roughly 30,000 agents run at once behind an online monitor that flags about 100,000 transcripts a week.
The score is a Claude judge grading Claude.
Every figure is self-measured, and any team can shard tasks smaller or relabel collaborates as leads while nothing real changes. If share of AI-led work becomes a standard line item by 2031, Goodhart's Law arrives with it.
2031: Every Team Publishes a Ladder Score
Pull back five years. Git went from a hacker tool to a hiring requirement in about a decade, and CI pipelines followed the same path. My read is that "share of AI-led work" rides the same arc. By 2031 it is a line item most knowledge teams report, whether or not the number means much.
That is the asymmetric bet. If you instrument your workflow now and the metric turns out to be noise, you lost a spreadsheet. If you skip it and the metric becomes the yardstick, you are five years behind on task hygiene. Headcount buys output. A task catalogue buys leverage.
Here is the warning half. Goodhart's Law says a measure that becomes a target stops being a good measure. Anthropic's 26%12 is a Claude judge grading Claude. Any team can shard tasks smaller or relabel "collaborates" as "leads," and the dashboard climbs while nothing real changes.
I cannot tell whether the 26%12 measures autonomy or just execution throughput. Anthropic says humans there still set goals, gate safety, and decide what ships. Claude does not pick research directions or reallocate compute. The metric tells you where the intern sits, not who runs the lab.
Two contrasts worth holding onto. Amateurs count how much the agent does; leaders count where the agent does it. A hundred agent-led log-parsing tickets carry less risk than one agent-led evaluation design. I think the Ownership Gap needs a second axis, stakes, before it means anything outside Anthropic's walls.
Impermanence cuts both ways here. Anthropic frames the index as a prototype it expects to revise. Dario Amodei's September 12 essay asked the industry to accept embedded third-party evaluators, and this post reads like the data those evaluators would check. An Anthropic worker publicly quit on September 9 over risk concerns, so the same lab produced both stories in the same month.
Score Your Own Task Catalogue by Friday
Show, don't tell. You do not need a frontier lab or a CS degree to run this exercise. You need a spreadsheet, one week of Slack history, and an honest hour.
First, inventory. Pull every recurring task your team did last week and aim for 30 to 50 line items. "Wrote the weekly metrics summary" counts. "Thought about strategy" does not.
Second, rate each one on the Epoch AI ladder. AL0 means no AI touched it. AL3 means an agent did big chunks under your direction, and AL4 means you gave a prompt and only reviewed the output. Be brutal, because the whole value is in the honesty.
Third, compute your own Ownership Gap. Take the percent at AL3 or higher and subtract the percent at AL4. Expect a wide gap. Assist is everywhere and ownership is rare, and that is normal on a first pass.
Fourth, pick three bucket-two tasks and rewrite them as tickets with a clear input, output, and check. Hand one to an agent and let it break, because it will break the first time. Log why it broke. That log is your monitor funnel in miniature.
Two tool notes from today's digest. ThursdAI's September 2026 ledger counts 50 AI releases through September 17, 11 of them last week. Read the ledger, not the launch posts, and only add a "new model" row to your sheet when a rating actually moves.
Same skepticism applies to tool directories. The Featured slot on the Rosebud AI listing is reportedly a paid Hola AI placement carrying a ref=featured tag, not a ranking. A dashboard number is only as good as who scored it, and that applies to Anthropic's 26%12 too.
Get your reps in. Rerun the sheet in four weeks. If your AL4 share moved, you learned what ticket shape your agents can own. If it did not move, you learned that too, and learning it is the whole point of measuring.
Score your own task catalogue before Friday.
- Inventory one week of real work. Pull every recurring task your team did last week from Slack and docs, and aim for 30 to 50 line items. "Wrote the weekly metrics summary" counts; "thought about strategy" does not.
- Rate each line on the Epoch AI ladder. AL0 means no AI touched it, AL3 means an agent did big chunks under your direction, AL4 means you gave a prompt and only reviewed output. Then subtract your AL4 share from your AL3-or-higher share to get your own Ownership Gap.
- Rewrite three assist-only tasks as tickets. Give each a clear input, output and check, hand one to an agent, and log exactly why it breaks the first time. Rerun the whole sheet in four weeks and see whether your AL4 share actually moved.
The 26%12 is a plumbing score, not an autonomy score.
Anthropic did not discover a model that builds models; it rebuilt its own research work until an agent could hold a ticket end to end. The 64-point gap between 90% participation and 26% ownership is the part any team can copy, because it is made of inventories, prompts, checks and monitors rather than frontier weights. The costs travel with it: safety compute doubles to 12% inside the AI-led slice, and about 50 flagged transcripts a week still land on a human. Treat the number as a diagnostic of task legibility, not a leaderboard, and remember it was graded by a Claude judge reading Claude's work. Score your own catalogue, then argue with the result in four weeks.
