Skip to main content
Back to News
news/AI Research

Claude Now Leads 26% of Anthropic AI Research Work

Anthropic says Claude now leads 26% of its AI research and development work, while more than 90% involves AI collaboration, raising the bar for oversight.

Stefan Trbojevic

Stefan Trbojevic

18 September 20262 min read
LinkedIn

The takeaway

Claude’s 26% share does not prove recursive self-improvement, but it shows that AI-led R&D is becoming measurable and requires operational oversight.

Why it matters for builders

Agent builders should track action coverage, review latency, escalation rates, and intervention paths, not only model benchmarks.

Claude Now Leads 26% of Anthropic AI Research Work

Anthropic says Claude now leads 26% of the company’s AI research and development work, a sharp change from less than 1% earlier this year. The result is not a claim of autonomous self-improvement, but it is a concrete sign that frontier labs are using AI to accelerate the work of building future AI.

What Anthropic measured

In a new measurement report from Anthropic, the company describes an R&D Automation Index based on an “Automation Level” scale developed by Epoch AI. At the level Anthropic calls “AI leads,” Claude can complete most of a task from a high-level prompt while a human supervises.

Anthropic reports that Claude leads 26% of its measured AI R&D work as of August 2026. More than 90% of the work is at or above the “AI collaborates” level, where the model completes large chunks under close human direction. The company also says Claude is not fully autonomous for any measured subset of the work.

Oversight becomes the real bottleneck

The same report says roughly 30,000 agents were doing research and engineering work on Anthropic’s most-used internal platform in August. Anthropic reports that online monitors covered 100% of those agents’ actions before execution, while offline systems ingested 100% of actions afterward. About 0.002% of more than one billion decisions were blocked by the online monitor.

Those figures are self-reported, and Anthropic acknowledges that cross-lab comparisons need common methods and independent verification. Still, the direction is important: the operational question is shifting from whether agents can contribute to research toward whether teams can measure, review, and intervene in that contribution quickly enough.

Why builders should care

For AI builders, this is a practical warning against treating agent capability as the only metric that matters. Production systems need action coverage, review latency, escalation rates, and clear intervention paths. A workflow that lets an agent write code, change infrastructure, or call external tools without those controls is not “autonomous” in a useful engineering sense. It is merely under-instrumented.

Key takeaway: Claude’s 26% share does not prove recursive self-improvement, but it shows that AI-led R&D is becoming measurable. Builders should instrument agent work with the same seriousness they apply to uptime, security, and cost.

Share𝕏

The Automation Brief

Read 5 AI stories instead of 50.

The essential moves in AI agents, models, automation and infrastructure — filtered for builders and operators, with the part that actually matters.

No noise. Unsubscribe anytime.

Editorial notes

Reported by

Stefan Trbojevic

Edited by

n8n Lab Editorial

Published

18 September 2026

Updated

18 September 2026

AI disclosure: AI assisted with research and drafting. Factual claims are reviewed by an editor.

n8n Lab is an independent service provider. We are not affiliated with, endorsed by, or sponsored by n8n GmbH. “n8n” is a trademark of n8n GmbH and is used here only to describe the platform-specific implementation and automation services we provide.