A pretraining researcher quits Anthropic over race-to-superintelligence dynamics; a current alignment researcher replies with a personal probability and an admission that should worry anyone tracking AI governance.
What Happened
Reported first by the Wall Street Journal late on September 8 (paywalled; exact headline and byline unverified) and confirmed by the men’s own posts on X, Jacob Coxon — a pretraining researcher who spent approximately three years at OpenAI and then Anthropic — announced he had resigned and is leaving the AI industry. He is not an executive, not a member of a safety team, and not a policy officer; his work was in pretraining. His opening post was direct: “I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.” He added that “accepting this race and entering the ‘endgame’ is a hubristic gamble that should not be launched from a private company’s Slack,” and told the WSJ: “we’re on track for a lot of the most aggressive of these scenarios where by the end of next year things could be out of control already.”
Roughly ninety minutes later, Evan Hubinger — a current Anthropic alignment researcher who leads the company’s alignment stress-testing work (the “alignment science lead” title appearing in some coverage is not confirmed) — quote-replied. Hubinger did not resign. His reply endorsed Coxon’s warning and then went further, in his own words: “Jacob is correct here — we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.” He continued — and this is an excerpt; his post extends past this point — “I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
Several details require precision before any structural reading. The >10% figure is explicitly personal — Hubinger’s own framing is “I personally think,” not a prediction and not Anthropic’s position. Coxon’s separate line that colleagues “believe it could kill us all by the end of the decade” is him characterizing a shared belief at the labs; it is a distinct claim from Hubinger’s personal probability, and blurring the two misrepresents both. This was one resignation followed by a colleague’s reply eighty-three minutes later — not a coordinated walkout, not a mass exodus. The name is Jacob Coxon, not Chris Coxon, as some early write-ups had it.
The key insight: Hubinger’s admission is not an accusation that Anthropic is reckless. It is something more structurally significant: a person whose job is stress-testing alignment at the most safety-branded frontier lab has said, in public and in the first person, that his employer is “trying its best” and “does not yet have a plan” — simultaneously. That combination is the crux. “Trying its best” and “no plan yet” in the same sentence is not a contradiction; it is an honest description of a capability-control gap that every lab faces and that almost none state this plainly.
The Structural Read
Strip the drama and what remains is a governance argument made from the inside, in the first person, with a number attached. That is a new register. Two days earlier, OpenAI chief scientist Jakub Pachocki made an institutional case — from a platform, in an essay — that no lab has solved alignment well enough to justify maximum scaling speed. Meta, in the same week, shipped a consumer agent it conceded could be prompt-injected. Those are external and corporate-level signals of a widening capability-versus-control gap. What Coxon and Hubinger added is qualitatively different: an individual researcher quitting and an individual alignment researcher replying, both speaking in the first person, one putting a personal double-digit probability on catastrophe, the other attaching his name to “we do not yet have a plan.”
The business consequence is not about the melodrama of the exchange. It is about social license. The implicit bargain under which private labs are permitted to build toward superintelligence rests on a claim: trust us, we are being responsible. Anthropic has made that claim more explicitly than any other frontier lab — safety is not a feature, it is the brand. When the person whose job is to stress-test that safety commitment says publicly that the lab “does not yet have a plan to solve alignment for superintelligence and is not clearly on track to” — while simultaneously affirming it is “trying its best” — the claim weakens. Not catastrophically, and not because anyone is lying, but because the honest accounting of the gap is now on the record, in the first person, from an insider.
That weakening is precisely the argument Pachocki made for external, enforced safety bars rather than voluntary self-governance. The Coxon-Hubinger exchange is its internal counterpart — and it lands at a moment when Anthropic is reported by multiple outlets to be marketing a listing valued in the trillions. That timing link is circumstantial and is the outlets’ framing, not the researchers’. But the tension between a public-market story built on being the responsible lab and an alignment researcher’s public “no plan yet” is structural, not circumstantial.
Evan Hubinger — Current Anthropic Alignment Researcher (excerpt; post continues)
“I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
Permission Layer — Business Engineer Framework
Social License Is the Invisible Governance Layer
The Permission Layer framework identifies the governance layer — regulation, social license, or both — that controls which AI capabilities can be deployed and at what speed. Anthropic’s social license is its primary differentiator from OpenAI and Google. When the people doing alignment stress-testing publicly price the downside in double digits and say there is no plan yet, the Permission Layer erodes — which is the concrete mechanism by which Pachocki’s case for external safety bars gains real-world traction. This is not about one resignation; it is about who gets to decide how fast the race runs.
Three Implications
IMPLICATION 1 — THE RESPONSIBLE-LAB BRAND IS NOW STRESS-TESTED IN PUBLIC
Anthropic’s competitive advantage over every other frontier lab is not its models — it is its claimed alignment leadership. Hubinger’s public statement is not an attack on that brand; it is an honest accounting from inside it. But “trying its best / no plan yet” is a harder investor, regulator, and talent story to tell than “safety first.” The brand absorbs this, but it cannot absorb a second or third similar disclosure at the same candor level without the social-license gap becoming a material governance question — especially if a public listing is live.
IMPLICATION 2 — PACHOCKI’S EXTERNAL-BARS ARGUMENT GETS INTERNAL CORROBORATION
When Jakub Pachocki argued on September 6 that no lab has solved alignment well enough to justify scaling at maximum speed, it read as one chief scientist’s institutional view. Hubinger’s admission — from inside the lab that most credibly claims to have prioritized alignment — now corroborates the structural claim from a different direction. The argument for externally enforced safety standards, rather than voluntary self-governance, is stronger today than it was 48 hours ago. Regulators reading both statements have more ammunition than they did last week.
IMPLICATION 3 — PERSONAL PROBABILITY IS A NEW UNIT OF AI GOVERNANCE DISCOURSE
The most durable thing Hubinger introduced is not the specific number — it is the unit. A named, senior insider voluntarily attaching a personal double-digit probability to catastrophe, in public, while still employed, is a new register for this conversation. It is harder to dismiss than an anonymous survey, a policy paper, or an executive’s prepared remarks. If this becomes a norm — researchers at safety-critical positions publicly calibrating and publishing their personal risk estimates — the governance conversation changes in structure, not just in temperature. Labs will need a position on whether that transparency is encouraged, discouraged, or somewhere in between.
The Bottom Line
One pretraining researcher resigned; one current alignment researcher replied. Neither is an executive, neither speaks for Anthropic, and this is not a coordinated exodus. What it is — after every caveat holds — is the first time a named insider at the industry’s self-styled safety leader has said, in the first person and in public, that the odds of catastrophe are above one in ten and that his own lab does not yet have a plan, while affirming it is trying its best. That combination — honest, specific, inside, and on the record — is the internal counterpart to every external pressure this week, and it is what makes this more than a resignation story. The social license that lets private companies race toward superintelligence is not a legal instrument; it is a running collective judgment about whether the people doing the work can be trusted to know what they are doing. That judgment got harder to sustain today. This article is not investment advice.
Sources: Jacob Coxon and Evan Hubinger posts on X (via @hilbertspaess, Sep 8–9, 2026); Wall Street Journal report, late Sep 8, 2026 (paywalled; exact headline and byline unverified, attributed); 91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
This is business analysis, not investment advice. Jacob Coxon is a pretraining researcher (not an executive or safety officer) who resigned; Evan Hubinger is a current Anthropic alignment researcher who has not resigned and whose “alignment science lead” title in some coverage is unconfirmed. His “>10% within the next decade” is an explicitly personal probability, not a prediction or an Anthropic position, and the “not clearly on track” line is an excerpt of a longer post. Quotes are as reported by the WSJ and the men’s own posts; Anthropic had not publicly responded as of publication.









