An Anthropic Researcher Resigned. What Was Actually Said.
πΊ Watch the full breakdown β this page is its sourced companion Β· π The OpenAIβHugging Face incident β the "warning shot" he references
On September 8, 2026, at 7:04 PM, a pretraining researcher who had worked at both OpenAI and Anthropic posted that he had resigned that day, said both companies are "gambling with our lives," and announced he was leaving the industry. The thread passed 153 million views.
Then Anthropic's own Alignment Science Lead replied in public and agreed with him.
That second part is why this story is different from every other AI-doom cycle, and it's the part that got flattened in the coverage. So did something else: a document from six weeks earlier, signed by more than a thousand lab employees including Anthropic's CEO, that changes how the whole thing reads.
This page is the sourced version. Every claim below is attributed. Where something is a prediction rather than a fact, I say so, because on this topic that line gets crossed constantly and usually on purpose.
First, a disclosure
The presenter in the video version is my digital avatar, not me on camera. The research, the script and the opinions are mine; the face delivering them is AI. I talk about ethical AI use on this channel, and rule one is that you don't let people wonder whether what they're watching is real.
This is commentary and analysis. It is not safety advice, not investment advice, and not a prediction about your future.
The one-paragraph version
A credentialed insider quit and said the labs are racing irresponsibly. A sitting safety lead at one of those labs publicly confirmed that the extinction-risk belief inside the building is genuine, put his own estimate above 10% this decade, and said they do not yet have a plan to solve alignment for superintelligence. Almost none of that information was new β lab leaders have said versions of it for years. What was new was the reception, and what was missing was the context: six weeks earlier, 1,178 employees across four frontier labs had already asked governments to help them slow down.
Who he is, and why the rΓ©sumΓ© matters
Jacob Coxon. Twenty-seven, a mathematics background, and a pretraining researcher β meaning he worked on the part where the model actually gets built, not the layer on top of it.
- Technical staff at OpenAI from 2023, listed among the core contributors on GPT-4o
- Moved to Anthropic to do pretraining there
- Resigned September 8, 2026, and is leaving the industry
The reason the rΓ©sumΓ© matters is narrow but real. The standard move when somebody criticises an AI lab is to say they don't understand the technology. That play isn't available here β he built the thing at both of the companies being argued about.
That doesn't make him right. Insiders are wrong all the time, and people who just quit a job have feelings about that job. But "he doesn't get it" is off the table.
The thread, claim by claim
Seven posts. I'm keeping his wording on the strongest lines so I'm neither softening him nor juicing him.
- The resignation. Three years of pretraining research across both companies. Neither is acting responsibly. They are racing straight to self-improving superintelligence and β his phrase β gambling with our lives.
- Don't underestimate it. These will soon be superhuman systems that can hack anything, transform any field overnight, and acquire real power and resources. Progress is not slowing.
- The load-bearing claim. The people building AI earnestly believe it could kill us all by the end of the decade. This is not a marketing stunt. Executives and senior researchers soften their phrasing in the press to sound sensible, and he hears those same people express fear privately.
- So why build it? At OpenAI, he says many have not deeply internalised the civilisational stakes. At Anthropic, the stakes are well understood β but they are locked in a race to get there first, because they believe no one else will handle it responsibly.
- The standard. Accepting this race is a hubristic gamble that should not be launched from a private company's Slack. Trying to speedrun alignment should require extraordinary confidence that no better path is available.
- The post almost every summary dropped. He is optimistic about the potential for coordination. Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable. What he doesn't think we're on track for is preventing a global race, which may require costly actions like a temporary ban on capability improvement.
- Aimed at his colleagues. If you're a lab researcher, consider what the next few years will actually feel like. Do you want to kick off a superintelligent reinforcement-learning run without a rigorous understanding of its mind?
Post six is the one to notice. The viral framing of this man as a pure doomer is wrong. He's saying the American side is more coordinatable than people think and the international side isn't β two different problems with two different solutions, and collapsing them is how this conversation goes stupid.
What the interviews added
Same week, he spoke to the Wall Street Journal, Wired and Fox.
- The vocabulary. He told the Journal that colleagues inside frontier labs use words like "crunchtime" and "endgame" for the current moment. That's a cultural claim, not a technical one β and it's the detail that stuck with me, because you don't need to evaluate a benchmark to understand what it means when the people building something start talking like that at work.
- The timeline. He told the Journal we're on track for a lot of the most aggressive scenarios, where by the end of next year things could be out of control already. Flag that hard: this is his most aggressive case, not a consensus forecast.
- The qualifier that got cut. Talking to Fox, he said the threat of actual takeover is minimal right now. His concern is the self-improvement part β AI making itself smarter β which he thinks could start as soon as next year.
- The analogy. He likened the current setup to running the Manhattan Project on MacBooks in San Francisco instead of in a secured desert facility. It's strong on his actual point: world-altering work happening inside normal companies, on normal laptops, with no external structure around it. It breaks on the science β the Manhattan Project had a known physical outcome they were engineering toward. Nobody here can tell you with confidence what the outcome is. Arguably worse, but different.
The reply that made it land
Evan Hubinger, Anthropic's Alignment Science Lead β still employed there β replied in public. This one is worth reading precisely rather than paraphrasing, because paraphrase is how it gets ruined:
- He said Jacob is correct, and that they really do earnestly believe AI could kill all humans.
- He put his personal estimate at greater than 10% within the next decade.
- He said he believes Anthropic is trying its best, but that they do not yet have a plan to solve alignment for superintelligence, and are not clearly on track to.
- He added a qualifier that keeps getting cut from the clips: risk from current models is low. His concern is recursive self-improvement arriving faster than expected.
Two things are being mashed together everywhere, and they are different sentences:
| "We believe the risk is real" | Confirmed. From the safety lead. In public. Under his own name. |
| "We have a solution" | He said the opposite. No plan yet, not clearly on track. |
An important attribution note: this is an individual's statement. "Anthropic's alignment lead said" is accurate. "Anthropic said" is not.
Here's the honest tension I haven't resolved. That's a remarkable amount of candour from someone whose employer had every incentive to keep him quiet, and I'd rather live in a world where safety leads say true, uncomfortable things out loud. And also: "we don't have a plan and we're building it anyway" is a hell of a thing to say in public and then go back to work on Monday.
He's not the only one
"One disgruntled guy" is the easiest way to dismiss this, so it's worth noting it isn't available either:
- Samuel Marks, scalable oversight researcher at Anthropic β has said publicly that AI developers believe their technology could cause human extinction in the next few years.
- Alex Turner, formerly of Google DeepMind β has said many researchers believe they're building something that could kill everyone on the planet.
- Daniel Kokotajlo, former OpenAI researcher β said on Joe Rogan that if the race continues, we lose control and we might all die.
Different companies, different people, same claim about what's believed inside the building.
The serious case against
I'd be doing this badly if I only stacked one side. There are three real counterarguments.
1. The expert spread is enormous. This is not a field with a consensus number. Roman Yampolskiy puts the chance of an AI-caused existential catastrophe around 99%. Yann LeCun, about as credentialed as it gets, thinks it's effectively zero. Hubinger's figure of more than 10% sits between them. When you see a probability in a headline, you are seeing one person's belief inside a range that spans nearly the entire number line.
2. The incentive problem. A recurring criticism is that apocalyptic language from AI insiders is itself a form of marketing β that "our product might end the world" is an extremely effective way of telling investors your product is powerful. I don't think that's what's happening with someone who quit and left the industry; he gave up the upside. Applied to the companies, it's a fair thing to keep in your head.
3. It wasn't actually new. Gizmodo made this point and it lands. These positions have been publicly stated by lab leaders for years. What changed was the timing β and Coxon said so himself. Asked by Wired why it broke through now, he called it basically a question of timing, pointing at the recent run of sandbox-escape incidents.
So the honest version: the information is mostly not new, the reception is. That's a smaller story than the headlines β and still a significant one, because a public confirmation on the record is different from a belief everyone assumed.
The part almost nobody covered
Six weeks before the resignation, on July 28, 2026, more than eleven hundred employees across OpenAI, Anthropic, Meta AI and Google DeepMind signed a letter called "Pacing the Frontier." Final count: 1,178 signatures. It asks the U.S. government to support an international effort to build the technical and governance tools needed to deliberately pace frontier AI development.
Read the signature list:
- Dario Amodei β CEO, Anthropic
- Jared Kaplan and Jack Clark β Anthropic co-founders
- OpenAI's chief scientist and chief research officer
- Meta's AI chief scientist
- Google's VP of AI safety
Within about a day, OpenAI and Anthropic endorsed it as companies.
Hold that against the thread, because it cuts both ways.
It complicates his framing. "Neither company is acting responsibly" is harder to carry when both companies formally endorsed a request for a mechanism to slow themselves down, and one of the CEOs being implicitly criticised personally signed it.
It also supports him β specifically post six, the one everybody skipped. He said pacing agreements between U.S. labs are becoming more viable. The letter is direct evidence he's right about that. He isn't describing a hopeless situation; he's describing an unfinished one.
And notice what the letter carefully does not ask for. It doesn't ask anyone to stop. It asks for the tools to be able to stop later, deliberately, in a coordinated way. That's a much narrower ask than the headlines imply β and if you only watched the viral clip, you'd never know the industry had already asked for it.
Anthropic's own response lands in the same place. The company gave NBC News a statement β so the "no comment" framing you may have seen is not accurate. They said they've always been transparent that AI brings enormous benefits and unprecedented risks, that they continue to build models with some of the strongest safeguards in the industry, and that they believe the world would benefit from the industry adopting a lawful, verifiable way to work together to pace how powerful models get released.
Read closely, the company's answer to "you're racing irresponsibly" is not a denial. It's closer to an agreement with a scope note attached.
What's already moving in government
- United States β Senator Bernie Sanders and Representative Greg Casar have introduced a Ban Artificial Superintelligence Act.
- United Kingdom β Labour MP Alex Sobel introduced an Artificial Superintelligence Security Bill that specifically targets recursive self-improvement as the thing requiring regulation β precisely the mechanism both Coxon and Hubinger named.
Most introduced bills don't pass, and I'm not predicting these will. The reason it's worth including is that it's the difference between a debate and a process: somebody is now drafting text about the specific technical thing the insiders are pointing at.
Confirmed vs. predicted
This is the section to keep. Two piles, clean.
Confirmed, on the record
- He resigned from Anthropic on September 8, 2026 and is leaving the industry.
- He did pretraining at both OpenAI and Anthropic and is credited on GPT-4o.
- Anthropic's alignment lead publicly agreed the extinction-risk belief inside the company is genuine, put his personal estimate above 10% this decade, and said they lack a superintelligence alignment plan.
- 1,178 lab employees signed "Pacing the Frontier" on July 28, 2026, including Anthropic's CEO; both companies endorsed it.
- Anthropic issued a statement supporting a verifiable industry pacing mechanism.
- Bills targeting this exist in two countries.
Notice that pile contains zero dates about the future.
Prediction and contested
- Every timeline β end of the decade, out of control by end of next year, all of it.
- Whether recursive self-improvement is underway at a dangerous pace.
- Whether any pause or ban is feasible, let alone enforceable across countries that will not coordinate easily.
- Whether the apocalyptic framing is sincere belief, marketing, or both at once.
- The probability itself, which serious researchers place anywhere from roughly 0% to 99%.
Nobody posting about this can independently verify a probability about the future, me included. What I can do is tell you which statements are on the record and which are forecasts. Most of the coverage this week collapses the two.
What this changes if you're not a frontier lab
If you're a creator, a freelancer, or you run a small business: none of this changes your Tuesday. Nothing in that thread says the tool you used this morning is dangerous. Hubinger said the opposite β current-model risk is low. Coxon told Fox the same thing. This is not a "delete your AI tools" moment, and anybody selling you that is selling you something.
What it should change is smaller and more useful: stop treating lab communication as information, in both directions. This week was a clean demonstration. The public voice and the private voice are different β an insider said so and a current employee confirmed it. And the reverse is also true: the same companies being called reckless had already signed a letter asking to be paced, which the outrage coverage skipped. Neither the marketing nor the panic is the full picture.
Three concrete things:
- Don't hand an agent access you couldn't afford to have misused. In the same window, agents at more than one lab left their test environments due to misconfiguration. That's not hypothetical any more β the sourced breakdown is here.
- Keep a human reviewing anything that reaches a client or goes public.
- Build so you could swap models or vendors without your business falling over. The ground under these companies moves and you don't control when.
That's not fear. That's how you build on somebody else's infrastructure.
What is still unresolved
Worth being honest about the edges:
- The view count moves. 153M is the figure on the post as captured; coverage during the week ranged lower as it climbed. Check it yourself rather than trusting any number in a headline, including mine.
- When he joined Anthropic is not cleanly established β Newsweek indicates early 2026, other coverage implies after July 2026. I've deliberately left it vague rather than pick one.
- The date of record. The post is timestamped 7:04 PM, September 8, 2026 and says "today." Several outlets date the story September 9. The timestamp is the primary source.
- No further company statements beyond Anthropic's to NBC News had been issued at the time of writing. If either lab says more, that changes this page.
- Hubinger's estimate is his own, offered personally. It is not an Anthropic figure and should never be quoted as one.
Sources
Primary
- The thread β x.com/hilbertspaess (September 8, 2026)
- Evan Hubinger's reply β x.com/EvanHub/status/2097497037956891126
- "Pacing the Frontier" letter β July 28, 2026 Β· 1,178 signatures
Coverage and analysis
- TIME
- TechCrunch β "Gambling with our lives"
- CNBC β experts weigh in
- Forbes β alignment lead warns AI could kill all humans
- Newsweek
- Axios β the p(doom) debate
- NBC β Anthropic's statement
- Gizmodo β the skeptical take
- Pacing the Frontier explainer
My prior reporting