Anthropic disclosed on September 17 that its Claude models are now "leading" 26% of the company's own AI research and development work, marking the first time a frontier lab has put a hard number on how much of its next-model pipeline the previous model is already driving.
What "Leading" Means Here
Anthropic defines a task as Claude-led when the model "can complete most of a task end-to-end from a high-level prompt, while a human supervises," and stresses that "Claude is not operating fully autonomously for any measured subset of AI R&D work." On the broader measure — tasks where AI performs "large chunks of work under close human direction" — the number climbs above 90%, meaning the model touches almost every research workflow in the building.
Built On Epoch AI's Scale
The metric is bolted onto Epoch AI's automation rating scale, an external framework for quantifying how much of a job an AI system is doing versus a human. Anthropic packaged it inside a three-part measurement framework: AI-led R&D automation levels (tracked since August 2025), agent oversight monitoring — activity surveillance, review timing and flagged behaviours — and compute allocation tracking that shows how much silicon Anthropic pointed at Claude-run research.
A Recursive Self-Improvement Story With Guardrails
CEO Dario Amodei has framed the disclosure as part of a wider push to publish safety-relevant metrics as AI systems start meaningfully accelerating their own successors. The company argues that measuring automation levels, agent oversight and compute share is now a prerequisite for reasoning about when to trigger safety interventions.
The disclosure lands as Anthropic scales up capacity and product surface area — from the $45B Nscale compute deal and Zerra DC Queensland anchor to Novo Nordisk's Claude Science drug-discovery deal — and gives regulators a first concrete peg for tracking how much of the next Claude is being built by the current one.
What It Doesn't Say
Anthropic did not break down the 26% figure by discipline (interpretability, RL, evals, infra), and it is a self-report rather than an independent audit — a caveat the company acknowledges. But it is a starting point, and Epoch AI-style scoring is easy for peers to adopt, which is likely the point.
Reporting based on coverage from Bloomberg, Engadget, The Washington Post and Anthropic.
