Anthropic, OpenAI, and the Staffing Problem at the Heart of AI Safety Auditing

Two frontier labs commit to independent evaluators with employee-like access. The binding constraint is not willingness — it is the only people qualified to do the job are the ones who just left the labs.

Sequence of Events — September 2026

Late August 2026

Joe Benton departs Anthropic’s safety team. No statement at the time.

11 September 2026

Benton publishes his reasons publicly on Substack. He names his next role: METR, an independent AI safety evaluator. His critique targets industry structure — he states explicitly it is “not about one particular company.”

12 September 2026 — Morning

Dario Amodei publishes a pacing essay committing Anthropic to give embedded external evaluators employee-like access. He cites METR as an example of the kind of body he has in mind: “Anthropic is unilaterally committing to this step now.”

12 September 2026 — Hours Later

Sam Altman states OpenAI will match the commitment: “Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We’ll have more to share soon.” Neither lab names a designated auditor, scope, or start date.

The sequence is stated as fact. Nothing here establishes that one event prompted another, and nothing suggests it did.

What Happened

On 11 September 2026, Joe Benton — who had left Anthropic’s safety team roughly two weeks earlier — published his reasons for going, and named where he was headed. Writing on Substack, he said: “I will soon be joining METR, an independent organization that evaluates how safe AI companies’ systems are.” He was careful about his framing from the first paragraph: his argument was about the industry, he said, and it is “not about one particular company.” Anthropic has not commented on his account.

His structural claim was direct. “AI capabilities are already improving extremely fast, and the companies are trying to go even faster,” he wrote, pointing to “increasingly concerning real-world incidents, including hundreds of agents from OpenAI recently hacking HuggingFace.” He noted that Anthropic “hasn’t had as severe an incident as the HuggingFace attack” — and attributed that partly to luck. That attribution is his characterisation and his opinion. His conclusion on the industry overall: “Humanity may not survive this transition.”

The following morning, Amodei published a pacing essay that committed Anthropic unilaterally to giving embedded external evaluators employee-like access, naming METR as an example — not an appointed auditor — of the kind of body he had in mind. Hours later, Altman matched it in a post, adding that OpenAI would have more to share soon. That is a stated intent, not a published policy. Neither company has defined scope, named a designated evaluator, or given a start date.

The key insight: Benton and Amodei share a diagnosis — competitive pressure structurally underinvests safety — and differ only on where the useful fix is located. One committed his company from the inside; the other moved to the body that does the checking. Both moves are responses to the same reading of the same incentive structure, and they landed within a day of each other.

Two labs have pledged employee-like access for outside evaluators. Neither has named a designated auditor, def
Two labs have pledged employee-like access for outside evaluators. Neither has named a designated auditor, defined the scope or given a start date — Amodei cites METR as an example, OpenAI has said only that more is coming. The tiers that would actually constrain capability remain at zero.

What Benton Did Not Say — and Why It Matters

It is worth being scrupulous about this, because the story is easy to misread. Benton does not accuse Anthropic of wrongdoing. He does not claim he was pushed out, and he does not allege knowledge of any undisclosed failure. Read as a whistleblower story, this is simply wrong. What he offers is a structural judgement about industry incentives from someone who worked inside one of the institutions those incentives govern.

His core claim — “Competition pushes every frontier company to underinvest in safety; the cost of falling behind is too high” — is a collective-action argument, not a company-specific allegation. That is precisely the mechanism Amodei’s essay is also built around, and precisely what his proposed industry coordination tier is designed to address. Amodei himself concedes the limits: “Some forms of coordination that would be impactful for pacing are legally challenging, and will require government support.”

The Structural Read

The binding constraint on the evaluator model is staffing, and the pipeline runs through the labs. Auditing a frontier AI system at the level of nuts and bolts requires people who understand how frontier AI systems are built. Overwhelmingly, those people have worked at a frontier lab. Benton’s move to METR is that pipeline operating exactly as it has to: safety talent flowing from the institution being evaluated to the institution doing the evaluation.

Every technical audit profession on earth began exactly this way. Accounting was staffed by practitioners from the firms it examined. Aviation safety boards drew from the airlines and manufacturers. Nuclear regulators were built from the engineers who designed the reactors. The path from insider to independent examiner is not a compromise of the model — it is the model, in its founding generation. What the profession then builds, over years, is the engineering that converts practitioner knowledge into durable independence.

BE Framework — Permission Layer

Engineered Independence vs. Inherited Independence

Independence in a technical audit setting cannot be inherited from an org chart — it has to be built. The components are four: publication rights that survive the subject’s objection; funding that does not come from the audited; tenure that does not depend on the relationship continuing; and recusal rules governing what a former employee may examine and when. None of this is a suggestion that METR is compromised — nothing reported suggests anything of the kind. It is a statement about what the commitments made this week now have to build, at scale, continuously, as the sector grows and the hiring pool remains the labs.

The independence question also scales non-linearly. If both commitments hold and the model spreads across the industry, the evaluator sector has to grow. Its hiring pool is the labs. Every hire is simultaneously a capability gain — someone who can actually read the system — and an independence question to be managed, not resolved once and filed away. This is not a solvable problem in the engineering sense; it is an ongoing governance problem of the kind that mature professions handle through published standards, mandatory disclosure, and rotation rules.

Dario Amodei — 12 September 2026

“Anthropic is unilaterally committing to this step now.”

The decisive test is already written into Amodei’s own formulation. Evaluators may redact material that is security-sensitive, legally privileged, commercially sensitive, or third-party confidential — but findings cannot be redacted merely for being unfavourable. That last clause is the whole thing. Whoever staffs the work, an evaluation that cannot publish an unwelcome conclusion is access without consequence.

Three Implications

IMPLICATION 1 — The Evaluator Sector Is Now a Strategic Constraint

Two of the three largest frontier labs have committed to employee-like evaluator access. Neither has named a designated auditor, defined scope, or set a date. The commitment is worth exactly what the evaluators can deliver — and the sector’s current capacity is unknown and unpublished. If the model spreads, evaluator scaling becomes a binding constraint on how fast the commitment can be honoured in practice, not just in language.

IMPLICATION 2 — Government Coordination Is the Remaining Variable

Amodei’s unilateral step is a first-mover signal, not a solution to the collective-action problem Benton identifies. Amodei says so directly: the forms of coordination that would most affect pacing are legally challenging and require government support. Voluntary evaluator access, however well-staffed, does not resolve competitive underinvestment in safety unless it is backed by requirements that apply uniformly. The policy gap between the commitment and the structural fix remains open.

IMPLICATION 3 — The Publication Clause Is the Audit’s Proof of Concept

The architecture of independence — publication rights, independent funding, tenure security, recusal rules — is not yet publicly specified by either lab or by METR in relation to these commitments. The clause that evaluators cannot redact merely unfavourable findings is the only publicly stated test. The first time an evaluator publishes a conclusion that the lab would have preferred to omit will be the moment the commitment proves its weight. Until then, it is a governance structure with one clause standing in for the rest.

Business Engineer Framework

The Permission Layer — How Governance Controls What AI Ships

The Permission Layer framework maps how regulatory and governance structures determine which AI capabilities reach deployment — and at what pace. The evaluator-access commitment made this week is a voluntary Permission Layer mechanism: labs granting external bodies the standing to slow or flag deployment. Whether it has teeth depends on the independence engineering underneath it. The full framework maps every governance layer from lab policy to international coordination.

Explore the Permission Layer Framework →

The Bottom Line

A safety researcher left a frontier lab and joined the body that checks frontier labs. The following day, two frontier labs committed to giving that kind of body employee-like access. Neither event caused the other, both are responses to the same structural diagnosis, and together they define the problem the industry now has to engineer its way through: the only people capable of auditing a frontier AI system are people who built one, which means independence cannot be assumed from the org chart — it has to be constructed, clause by clause, hire by hire, publication by publication. The commitment made this week is real. So is the distance between the commitment and the thing it has to become.


Sources: Joe Benton, “Why I Left Anthropic’s Safety Team,” Substack, 11 September 2026. Dario Amodei pacing essay, Anthropic, 12 September 2026. Sam Altman public statement, 12 September 2026. Anthropic has not commented on Benton’s account. Anthropic and OpenAI are private companies; Anthropic is reportedly preparing a public listing — no timing or valuation is expressed here. METR is a non-profit evaluation organisation; no figure for its headcount, funding, or budget is stated or estimated in this piece. This article is business analysis only and does not constitute investment

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

This account of Joe Benton’s departure and his stated plans comes from his own public post of 11 September 2026, with press amplification. Anthropic has not commented, and no independent verification of his account is asserted here. Benton criticises the industry generally and states explicitly that his argument is “not about one particular company.” He does not accuse Anthropic of wrongdoing, does not say he was dismissed or pushed out, and does not claim knowledge of any undisclosed failure; this is not a whistleblower disclosure. His remark that Anthropic’s comparatively better incident record owes partly to luck is his own characterisation and opinion. Nothing here asserts or implies that METR is captured, conflicted or insufficiently independent; no evidence suggests that, and the independence point raised is structural and prospective. No figures for METR’s headcount, funding or budget are known or implied. Benton published on 11 September and Dario Amodei’s essay appeared on 12 September. That is a sequence, not a causal relationship, and nothing indicates one prompted the other. Amodei’s essay cites METR as an example of the kind of evaluator he means rather than as an appointed auditor; neither Anthropic nor OpenAI has named a designated evaluator, defined the scope of access, or given a start date, and OpenAI’s position is a stated intent rather than a published policy. Anthropic and OpenAI are private companies, Anthropic reportedly preparing a listing; METR is a non-profit evaluation organisation. This is business analysis, not investment advice, no view is expressed on any security, and nothing here predicts further departures.

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading