Anthropic has proposed three public measures for AI-assisted research, agent oversight, and safety compute. The disclosures are unusually specific, but every substantive figure remains a company-reported result that outsiders have not reproduced.
The headline number is 26%. Anthropic says Claude now “leads” that share of its measured AI research and development work, up from less than 1% in February 2026. Yet “leads” does not mean autonomous—and the number does not establish that Claude is building its successor, doubling research output, or reducing costs.
That tension is the point of Anthropic’s September 17 proposal. It identifies three measurement families: AI involvement in R&D, agent-monitor operations, and AI-R&D compute allocation. The framework turns an abstract debate into quantities that developers, buyers, and policymakers can question.
But these are still Anthropic’s internal measurements. The underlying records, classifiers, monitoring recall, and task-level judgments were not independently inspected in the cited material. Until another party can reproduce them, the numbers should be read as structured vendor disclosures—not neutral industry benchmarks.
What Anthropic’s 26% Actually Measures
Anthropic’s prototype R&D Automation Index organizes research work on an AL0-to-AL5 scale. At AL3, Claude “collaborates.” At AL4, or “leads,” Claude can complete most of a task end-to-end from a high-level prompt while a human supervises. AL5 would mean fully autonomous operation with no human in the loop.
As of August 2026, Anthropic reports that 26% of measured work was AL4 and more than 90% was at least AL3. No measured subset reached AL5. That final detail matters: the company’s own scale distinguishes supervised task leadership from autonomy.
The index is elaborate, but not cleanly objective. For each week in July, Anthropic sampled 20% of staff in departments participating in the model-R&D loop. A Claude research agent produced roughly 15,000 granular tasks in a frozen tree with 542 nodes and 378 leaves. A separate Claude judge assigned automation levels. Anthropic weighted tasks using sampled person-time and calls the method a “crude approximation.”
The reported judge check illustrates why borderline classifications matter. Exact model-to-human agreement was 59%, while exact human-to-human agreement was 35%. Model and human scores landed within one level 97% of the time. That may indicate broadly similar ordering, but a one-level difference around AL3 and AL4 can directly change the headline share of work that Claude supposedly “leads.”
| Reported measure | What it can show | What it cannot establish |
|---|---|---|
| 26% of measured work at AL4 | Anthropic’s estimate of supervised, end-to-end task leadership | Research output, model quality, cost savings, causal acceleration, or autonomy |
| 100% stated monitor routing | Claimed pre-execution coverage and post-execution ingestion on one internal platform | Monitor precision, recall, false negatives, or coverage of every deployment |
| 6% and 12% safety-compute shares | Classifier-defined allocation during one week | Safety effort, staff attention, quality, effectiveness, or a trend |
Two Checks Beyond Task Automation
Agent Monitoring: Coverage Is Not Detection Quality
Anthropic says approximately 30,000 research and engineering agents were active at any moment in August on its most-used internal platform. It reports 100% online-monitor coverage before execution and 100% offline-monitor ingestion afterward. The online layer can block or redirect actions; the offline layer reviews activity after the fact.
Across more than one billion agent decisions that month, Anthropic says the online monitor blocked 0.002%, or about one in 47,000. Offline systems reportedly flag around 100,000 transcripts a week for automated review, with roughly 50 escalated to humans. Anthropic also reports that one to two transcripts per 1,000 are flagged for further review and that humans review blocked actions within one week.

Those figures describe routing and intervention frequency, not detection performance. The page provides no independently audited evaluation, known-misbehavior catch rate, or representative assessment of all Anthropic agent deployments. A low block rate could indicate low-risk work, conservative workloads, or narrow triggers. It cannot, by itself, prove that harmful behavior is rare or that the monitors reliably catch it.
Enterprise buyers should separate four questions: Does every relevant action pass through a monitor? How quickly does review occur? What gets escalated? How often does it miss known bad behavior? Anthropic offers detail on the first three within the disclosed platform. The fourth remains unanswered.
Compute Allocation Adds Context, Not a Safety Score
Anthropic’s third measurement is a snapshot from July 13–20. The company reports that about 6% of AI-R&D compute was allocated to safety, while about 12% of compute used for AI-driven AI R&D went to safety.
These are shares of fungible-capacity usage during one week, not fixed budgets or expense allocations. Anthropic classified workloads with a prompted Claude classifier. From almost 10,000 research training and evaluation runs, it sampled about 14%, weighting the sample toward high-compute runs. The labels were “best-effort,” safety and capabilities work can overlap, and safeguards-classifier compute was excluded.
The result may help track how a lab categorizes resource use, but compute is not a proxy for safety quality. A high share would not demonstrate that safeguards work; a low share would not capture labor or effectiveness. One week also cannot show direction over time.
The Best Counterargument: Imperfect Disclosure Beats No Disclosure
Dismissing the framework because Anthropic uses its own data and models would miss its contribution. The company publishes definitions, exposes judging disagreement, limits its agent-platform claim, acknowledges no common cross-lab method, and proposes third-party or cross-developer verification.
Anthropic’s August Risk Report supplies an important check on overheated interpretations: it says internal AI R&D is significantly faster with AI assistance, but not yet twice as fast, and that measurement remains difficult. It also says the models had not met either automated-R&D criterion in the company’s Responsible Scaling Policy as of July 15. Anthropic separately states that recursive self-improvement has not been reached.
So the right response is neither to accept the numbers at face value nor to ignore them. The disclosure creates testable questions. Can an outside evaluator reconstruct the taxonomy and weights? Does AL4 correlate with validated output or time-to-result? Can auditors examine monitor precision, recall, false negatives, and adversarial performance? Will disagreements and exceptions be published, rather than hidden behind an attestation?
What Decision-Makers Should Take Away
- Anthropic reports substantial AI participation in its R&D, but its 26% AL4 figure describes supervised task leadership, not autonomous research or proven recursive improvement.
- Claimed 100% monitoring coverage is operationally relevant, yet it says nothing by itself about missed harmful behavior.
- The 6% and 12% compute figures are one-week, classifier-defined allocation snapshots—not measures of safety effectiveness.
- All substantive results remain Anthropic self-reports. Comparable methods and independent access are the missing pieces.
For procurement and policy, the framework is best treated as a disclosure template. Ask vendors to define the unit of work, state who or what assigns labels, publish disagreement rates, bound platform coverage, and distinguish resource allocation from outcomes. The largest percentage is less important than whether someone outside the lab can audit how it was produced.
