Anthropic said Claude “leads” 26% of its AI research and development work as of August 2026, up from under 1% in February 2026 on an automation scale developed by Epoch AI. The measurement offers a view of how close a frontier lab may be to AI building its successor as developers debate whether to slow that process.
What Changed
- Anthropic said Claude "leads" 26% of its AI research and development work as of August 2026, up from under 1% in February 2026, on an automation scale developed by Epoch AI.
- The ratings were produced largely by Claude. Model and staff ratings matched exactly 59% of the time, while pairs of employees matched only 35% of the time.
- About 30,000 agents worked at once on Anthropic's most-used internal platform in August, and its online monitor blocked about one in 47,000 of more than one billion decisions.
- In one July week, about 6% of Anthropic's AI R&D compute went to safety work, rising to about 12% for AI-driven AI R&D.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
How the index works
The Anthropic R&D Automation Index rates tasks from AL0, with no AI involvement, to AL5, where AI works without a human in the loop. AL3 means Claude completes large portions under close direction. At AL4, “leads,” Claude can complete most of a task end-to-end from a high-level prompt while a human supervises.
For a failed data pipeline, an AL4 Claude could inspect logs, find and test a fix, handle surprises and compare the repaired output with the last good run. An engineer would still decide whether to deploy it. “Claude is not operating fully autonomously for any measured subset of AI R&D work,” Anthropic said.
For each week of July 2026, Anthropic randomly sampled 20% of staff in the departments that make up its model R&D loop. A Claude research agent reviewed each sampled person’s week using Slack and internal documentation, producing about 15,000 granular tasks. Claude arranged them into 542 nodes, and Anthropic froze that basket for later measurements.
FREE · ABOUT FIVE MINUTES
Track how fast AI labs are building AI.
Implicator cuts the flood to the developments worth your attention, with concise reporting on what changed and why it matters. Delivered every weekday at 4:45 a.m. Pacific, 7:45 a.m. Eastern.
From San Francisco. No spam. Unsubscribe anytime.
Who graded the grading
A Claude research agent assessed each category, then a separate Claude judge assigned an automation level. Anthropic employees rated their own work areas without seeing the model’s evidence. The model and employees matched exactly 59% of the time, while pairs of employees matched exactly 35% of the time. Model and human ratings fell within one level of each other 97% of the time.
The report does not say any outside party has checked the figures, which are Anthropic’s own measurements, largely produced by Claude. Anthropic says it plans to embed third-party evaluators and is setting them up. Anthropic also acknowledged that “the ‘judge’ model could make the same kinds of errors as the model it is checking.”
Agent oversight
About 30,000 agents were working simultaneously on Anthropic’s most-used internal research and engineering platform in August 2026. The figures cover that platform only. Every action passed through an online monitor, typically within seconds.
Of more than one billion agent decisions during August 2026, the monitor blocked about one in 47,000. Humans review blocked actions within one week. A separate offline system flags about 100,000 transcripts each week. Anthropic separately gave the flag rate as roughly one or two transcripts in every thousand. About 50 high-priority cases reach people weekly. Anthropic wrote, “We have been doing this kind of monitoring for months, not years.”
Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.
Safety compute and OpenAI
From July 13 through July 20, 2026, Anthropic classified about 6% of its AI R&D compute as safety work. The share was about 12% for compute used in AI-driven AI R&D. Anthropic called both estimates conservative because mixed-purpose work was counted as capabilities research, while safeguards classifiers were excluded.
Anthropic said one week is enough to show that the measurement can be made, but not enough to show a meaningful trend. Compute also measures spending rather than effort: a more efficient safety classifier can reduce safety’s compute share without reducing the work performed.
OpenAI said on September 6, 2026, that it had reached its goal of an “automated research intern.” During the preceding six months, more than half of its successful tasks estimated to take a person four to eight hours required at least one human intervention. The labs measure different activities with different methods, so their figures cannot be compared directly. Anthropic said it will periodically rebuild its task basket and continue publishing the results.
Frequently Asked Questions
What does it mean that Claude "leads" 26% of Anthropic's R&D?
On the Epoch AI automation scale Anthropic used, "leads" is level AL4: Claude can complete most of a task end-to-end from a high-level prompt while a human supervises. Anthropic said Claude reached that level for 26% of its AI R&D work as of August 2026, and is not operating fully autonomously on any measured part of the work.
How did Anthropic build the R&D Automation Index?
For each week of July 2026, Anthropic sampled 20% of staff in its model R&D departments. A Claude research agent reviewed their weeks using Slack and internal documents, producing about 15,000 tasks, which Claude arranged into 542 nodes. A separate Claude judge then assigned each category an automation level.
Has anyone outside Anthropic checked the figures?
The report does not say any outside party has checked them. They are Anthropic's own measurements, largely produced by Claude. Anthropic says it plans to embed third-party evaluators and is setting them up.
How much of Anthropic's compute goes to safety work?
In the week of July 13 to 20, 2026, Anthropic classified about 6% of its AI R&D compute as safety work, and about 12% of compute used in AI-driven AI R&D. It called both estimates conservative and said one week is not enough to show a meaningful trend.
How does this compare with OpenAI?
OpenAI said on September 6, 2026, that it had reached its goal of an "automated research intern," and that more than half of its successful four-to-eight-hour tasks in the preceding six months needed at least one human intervention. The two labs measure different activities with different methods, so the figures cannot be compared directly.
AI-generated summary, reviewed by an editor. More on our AI guidelines.



IMPLICATOR