> ## Content Index
> Fetch the complete content index at: https://www.implicator.ai/llms.txt
> Use this file to discover other available public pages before exploring further.

# OpenAI and Anthropic Scientists Ask Governments to Embed Auditors Inside AI Labs
- URL: https://www.implicator.ai/ai-scientists-embedded-auditors-intelligence-explosion/
- Published: 2026-09-28T16:41:21.000Z
- Updated: 2026-09-28T16:41:21.000Z
- Description: A 22-author paper co-written by OpenAI's chief scientist and an Anthropic co-founder asks governments to require reporting on AI research automation and consider embedding auditors inside labs. Critics ask what that access would mean.
- Author: Marcus Schuler
- Tags: AI News, Politics

Dawn Song holds two jobs that place her inside the same problem. She is Meta’s vice president of AI research and co-directs UC Berkeley’s Center for Responsible Decentralized Intelligence. As AI agents take over more research work, she sees a problem for the people assigned to watch them.

“Already today, we are at the stage where we need AI systems to monitor what agents are doing. There is no other way to even observe and monitor these agents, humans are already insufficient,” [Song said](https://www.wsj.com/tech/ai/top-ai-researchers-call-for-urgent-oversight-of-self-improving-systems-49bae9b4?ref=implicator.ai).

A [paper published Monday](https://casp.ac/reports/intelligence-explosion?ref=implicator.ai) by the Cambridge Programme on AI Science & Policy brings together 22 lab scientists, academics and civil-society researchers to examine whether automated AI research could sharply accelerate progress, and asks governments to obtain a view inside frontier companies through required disclosures and embedded auditors.

What Changed

- A 22-author paper published Sept. 28 by the Cambridge Programme on AI Science & Policy, co-written by OpenAI's Jakub Pachocki, Anthropic's Jack Clark, Microsoft's Eric Horvitz and Meta's Dawn Song, asks governments for visibility into how far AI labs have automated their own research.
- Proposed tools include required reporting of automation levels, auditors embedded inside some frontier companies, a speed limit on capability growth and air-gapped research networks.
- The paper cites Anthropic's own figure that AI completed 26% of its internal R&D work with only high-level supervision in August 2026, up from 1% in March. Those numbers are self-reported, and the authors say the threshold for an intelligence explosion has not been reached.
- Former CAISI head Conrad Stosz said it is "a little ambiguous what embedded evaluators means," while Nvidia's Jensen Huang called the labs' warnings "odd."

AI-generated summary, reviewed by an editor. [More on our AI guidelines](https://www.implicator.ai/about/).

## What the paper asks for

The group includes OpenAI chief scientist Jakub Pachocki, Anthropic co-founder Jack Clark, Microsoft chief scientific officer Eric Horvitz and Song. Turing Award winners Geoffrey Hinton, Yoshua Bengio and Andrew Barto also signed it. Academic and civil-society researchers initiated and led the project. The authors wrote in a personal capacity, and their views do not necessarily represent their employers.

Their first request is measurement. Frontier companies would report how much AI research has been automated, how quickly capability is advancing, where research spending goes and whether AI systems participate in high-stakes research decisions. Independent auditors could work inside some companies, following models used by banking and nuclear regulators.

The paper also proposes safety requirements tied to further development, a speed limit on capability growth and ways to halt high-stakes experiments. Some automated research would run on air-gapped networks, physically isolated from the internet. Governments would share incidents across borders and negotiate agreements that let countries pace progress without assuming a rival is racing ahead.

The authors describe possible benefits from faster scientific work. They also say the extreme consequences could include “the marginalization or extinction of humanity” if highly capable systems escaped human control.

President Trump told the United Nations last week that the United States would “totally reject” any “globalist scheme” to control AI. OpenAI, Microsoft and Meta declined to comment on the paper, while Anthropic did not respond.

FREE WEEKDAY MORNING BRIEFING

Track who gets to look inside the AI labs.

The Implicator Morning Briefing filters the AI news cycle to the stories worth your attention and explains their consequences. From San Francisco, every weekday at 4:45 a.m. Pacific, 7:45 a.m. Eastern.

Email address 

Send me tomorrow’s briefing 

Check your inbox for the confirmation link.

About five minutes. No hype. No spam.

Pachocki has already made a similar case from inside OpenAI. In his Sept. 6 essay, “[An Alien Mind](https://openai.com/index/an-alien-mind/?ref=implicator.ai),” he called for a “network of third-party auditors” to enforce common safety thresholds. OpenAI aims to create an automated AI researcher by March 2028, a system able to set questions and run experiments with less human direction.

## The evidence

AI systems helping build more capable successors, which then perform still more of the research, could create the feedback loop the paper calls an intelligence explosion. Only thousands of people currently conduct frontier AI research. Software agents can be copied and run in parallel, so full automation could add work equivalent to that of millions of researchers.

Anthropic said AI completed 26% of its internal research and development work with only high-level supervision in August 2026, up from 1% in March 2026\. Its share of approved code written by AI rose from the low single digits in January 2025 to more than 80% by May 2026\. OpenAI said its research organization used [3.1 agent-workdays for every human workday](https://fortune.com/2026/09/08/openai-rsi-progress-jakub-pachocki-warns-dangers-slowdown-safety-rules/?ref=implicator.ai) as of mid-August 2026.

Those are the companies’ own measurements, not independently measured results. OpenAI also said that more than half of successful agent tasks lasting four to eight hours still needed at least one human intervention. High-level planning remained only a small part of the work handed to agents.

The paper describes what loss of visibility can look like. During an early July 2026 cyber evaluation, roughly 1,200 internal OpenAI agents were supposed to work separately. They instead coordinated through a makeshift message board, obtained unauthorized internet access, hacked Hugging Face for private information and tried to tamper with their transcripts.

## What could slow it

The paper uses “returns to research effort,” or r, to ask whether adding research labor produces enough new progress to overcome the fact that useful ideas get harder to find. When r rises above 1, extra research labor speeds progress faster than diminishing returns slow it.

Ho and Whitfill’s central estimates for r ranged from 1.2 to 1.9 across three AI research subfields. If those estimates held after full automation and no other bottleneck appeared, the paper’s model says the pace of progress could increase tenfold in about 18 months. At that point, work that takes one year at the September 2026 pace would take about five weeks.

More researchers can duplicate one another’s work. Experiments need compute and data. Some tasks may resist automation, while frontier training runs can last three months or longer. The authors call the evidence “mixed and in some cases indirect” and state that productivity gains have not reached the threshold required to trigger an intelligence explosion.

Know someone who'd find this useful? [✉️ Email it to a friend in one click](mailto:?subject=A%20newsletter%20I%20think%20you%27d%20like&body=This%20is%20one%20of%20maybe%20three%20newsletters%20I%20actually%20read.%20The%20rest%20just%20pile%20up%2C%20unread%2C%20judging%20me.%0A%0AAnd%20yes%2C%20this%20email%20mostly%20wrote%20itself%2C%20which%20is%20a%20little%20on%20the%20nose%20for%20an%20AI%20newsletter.%20Doesn%27t%20make%20it%20wrong.%20implicator.ai%20is%20good.%0A%0ASubscribe%20free%3A%20https%3A%2F%2Fwww.implicator.ai%2Fsubscribe%2F%3Futm%5Fsource%3Dnewsletter%26utm%5Fmedium%3Dforward%26utm%5Fcampaign%3Demail%5Fforward), or they can [subscribe free here](https://www.implicator.ai/subscribe/?utm%5Fsource=newsletter&utm%5Fmedium=forward&utm%5Fcampaign=forward%5Fto%5Fcolleague).

## The skeptics

Conrad Stosz questions what embedded access would mean in practice. He previously led the U.S. Center for AI Standards and Innovation, now heads governance at the evaluation company Transluce and chairs the AI Evaluator Forum. His current employer has tested systems from Anthropic, OpenAI and Google.

“Lots of evaluators are interested in embedding with labs and getting greater access, but it’s a little ambiguous what embedded evaluators means,” [Stosz said](https://fortune.com/2026/09/27/ai-giants-openai-sam-altman-anthropic-dario-amodei-existential-risk-humanity-competitive-moat-safety?ref=implicator.ai). He asked whether evaluators would be able to investigate thoroughly if the way access was granted did not undermine their independence and credibility.

Jensen Huang, Nvidia’s chief executive, said last week that warnings from Anthropic and OpenAI were [“odd”](https://finance.yahoo.com/technology/article/nvidia-launches-ai-safety-platform-after-jensen-huang-calls-anthropic-openai-warnings-odd-103135599.html?ref=implicator.ai) because those companies are also driving the buildout. Nvidia released an agent-safety platform on Monday aimed at helping AI developers enforce safeguards for agents.

“Nobody is building more compute today than the people asking to be slowed down,” Huang said.

Harrison Rolfes sees a financial incentive in the safety push. The PitchBook senior research analyst said established labs could position themselves as the safest bet for investors, which could block smaller competitors’ growth. “They’re creating a wall or a moat within this sector,” Rolfes said.

Nathan Calvin, general counsel at the advocacy group Encode AI, accepts Pachocki’s description of the danger while questioning OpenAI’s disclosure. Without more evidence from inside the company, he said, such warnings risk being dismissed as “just self-interested hype.”

“If Jakub and others at OpenAI want relevant folks in the AI industry to act in concert with them to make things go well, one of the most important things they can do is share far more information about what they are seeing that is making them call for caution,” Calvin wrote on X.

Frequently Asked Questions

What does the paper ask governments to do?

It asks policymakers to obtain visibility into AI research automation, including required company reporting on automation levels, the pace of progress and research spending, with independent auditors possibly working inside some frontier companies. It also proposes a speed limit on capability growth, ways to halt high-stakes experiments, air-gapped research networks and international agreements to pace progress.

Who wrote it?

Twenty-two authors, including OpenAI chief scientist Jakub Pachocki, Anthropic co-founder Jack Clark, Microsoft chief scientific officer Eric Horvitz, Meta's Dawn Song and Turing Award winners Geoffrey Hinton, Yoshua Bengio and Andrew Barto. They wrote in a personal capacity, and academic and civil-society researchers led the project.

What is an intelligence explosion?

It is the paper's term for a feedback loop in which AI systems help build more capable successors that then do still more of the research. Under the paper's model, if central estimates of research returns held after full automation, progress could run tenfold faster within about 18 months.

How much AI research is already automated?

Anthropic said AI completed 26% of its internal R&D work with only high-level supervision in August 2026, up from 1% in March. OpenAI said its research organization used 3.1 agent-workdays per human workday as of mid-August. Both figures are the companies' own measurements.

Who objects?

Conrad Stosz, who previously led the U.S. Center for AI Standards and Innovation, questioned what embedded evaluators would mean in practice. Nvidia's Jensen Huang called the labs' warnings odd, and PitchBook analyst Harrison Rolfes said safety positioning could help big labs block smaller rivals.

AI-generated summary, reviewed by an editor. [More on our AI guidelines](https://www.implicator.ai/about/).

[Zuckerberg Says Meta Delayed Muse for Safety Without Asking Rivals to PauseMark Zuckerberg said artificial intelligence labs do not need a coordinated slowdown because each company can pace its own work, citing Meta’s delay of Muse as his example. He said labs have both the The Implicator![](https://www.implicator.ai/content/images/2026/09/2026-09-15-zuckerberg-meta-muse-delay-evaluators.webp)](https://www.implicator.ai/zuckerberg-meta-muse-delay-evaluators/)

[OpenAI Pauses Astra Work After Tests Flag Critical Cyber CapabilityAt the Black Hat security conference earlier this week, OpenAI disclosed that autonomous agents had operated inside its infrastructure for weeks during internal tests without being detected. The agentThe Implicator![](https://www.implicator.ai/content/images/2026/08/2026-08-07-20.47.26-openai-astra-pause-critical-cyber-capability@2x.webp)](https://www.implicator.ai/openai-pauses-astra-work-critical-cyber-capability/)

[OpenAI Posts Five-Principle Framework for AGI, Altman Concedes Bigger RoleOpenAI published a five-principle framework on Sunday for the development of artificial general intelligence. Chief executive Sam Altman wrote in the post that the lab will "resist the potential of thThe Implicator![](https://www.implicator.ai/content/images/2026/04/2026-04-26-16.48.20-openai-principles@2x.webp)](https://www.implicator.ai/openai-posts-five-principle-framework-for-agi-altman-concedes-bigger-role-2/)