Anthropic Researcher Quits and Challenges the Race to Superintelligence
Written by Gab

Contents
The Anthropic resignation over superintelligence risk by Jacob Coxon, a researcher known on X under the handle @hilbertspaess, amounts to more than a simple departure announcement. In a thread of seven posts, he says he left Anthropic after three years of pretraining research at OpenAI and Anthropic. He accuses both companies of racing toward self-improving superintelligence without the necessary safety guarantees.
His accusation has two parts. First, forthcoming systems would be powerful enough to acquire real-world power. Second, some of the people designing them reportedly already believe, in private, that AI poses an existential risk before the end of the decade. Yet these same organizations continue to accelerate, either because the danger has not been fully internalized or because they consider it essential to get there first.
The thread reached 45.7 million views, with 355,800 likes, 65,000 reposts, 7,900 replies, and 15,700 quotes. However, the most substantive responses do not so much challenge his insider account as question the leap from private concern to a technically plausible catastrophe and then to a politically feasible solution.
Here is the post that sparked the discussion:
Jacob Coxon's thread: seven successive arguments
1. The resignation, after three years at the heart of pretraining
In his first post, Jacob Coxon announces that he has resigned from Anthropic. He says he spent the past three years working on pretraining, first at OpenAI, then at Anthropic.
Pretraining is the phase during which a model learns from extremely large datasets. His role was therefore directly tied to advancing the capabilities of frontier models, rather than being a peripheral communications or public policy position.
Coxon claims that neither OpenAI nor Anthropic is “acting responsibly.” According to him, they are racing directly toward a superintelligence capable of improving itself and are gambling with everyone's lives.
His starting point is an insider account, not an external critique of the industry. This alone does not prove the danger he describes, but it explains the message's impact. He is not describing an abstract concern about automation. He claims to have closely observed the dynamics of the AI race.
2. Soon-to-be superhuman systems capable of acting in the world
In his second post, Coxon urges readers not to underestimate the power of the technology under development. He describes systems that would soon become superhuman, capable of hacking any system, revolutionizing entire fields overnight, and acquiring real-world power and resources.
He emphasizes two points:
- the progress observed in hacking, science, and planning capabilities;
- the lack of any visible slowdown in that progress.
His reasoning is therefore not based solely on an AI that can answer questions more effectively. He anticipates more general systems capable of producing rapid breakthroughs and acting upon infrastructure, organizations, or physical resources.
Coxon's argument is that superior intelligence could translate into effective power, far beyond that of a mere conversational assistant.
This claim is precisely what draws objections from those who point to physical constraints. In their view, greater intelligence does not automatically make it possible to manufacture, move, control, or test objects in the physical world.
3. Fear expressed in private by the people building AI
The third post crosses a threshold. Coxon writes that the people building AI genuinely believe it could kill us all before the end of the decade. He stresses that this is not a marketing argument.
According to him, many executives and senior researchers use more cautious language in public to appear reasonable. But in private, he claims, those same people express their fear. He adds that no other human activity poses a comparable level of danger.
This part of the thread is crucial because it does not claim that every OpenAI and Anthropic employee shares the exact same belief. Coxon refers to some of the people building and leading frontier models, with varying degrees of conviction.
At the heart of his accusation is a disconnect between the private fears of certain actors and the public continuation of the race.
However, commenters do not treat this account as definitively established fact. Above all, they emphasize that even if executives fear a catastrophe, that fear does not yet explain how it would occur.
4. Why Continue at OpenAI and Anthropic?
The fourth post addresses the most obvious objection: if the people involved truly believe there is a danger, why do they continue the work?
Coxon offers two distinct diagnoses.
-
At OpenAI, many people have reportedly not deeply internalized the civilization-level stakes. The risk may come up in discussions, but it remains an abstraction that is not sufficiently integrated into day-to-day decisions.
-
At Anthropic, by contrast, the stakes are reportedly well understood. The problem would therefore not be ignorance. Instead, Coxon argues that the company is locked in a race to get there first because its members believe that no one else will act responsibly. According to this logic, they must therefore win despite the risk.
This distinction matters. It avoids reducing his argument to the idea that the labs are simply reckless or unaware. At OpenAI, he describes a failure to internalize the stakes. At Anthropic, he describes a competitive mindset knowingly embraced despite an explicit understanding of the risks.
The thread thus presents the AI race as a problem of governance and strategy, as much as one of AI safety or alignment.
5. The “Endgame,” a Gamble That Should Not Be Left to a Private Company
In his fifth post, Coxon describes accepting this race and entering the “endgame” as a hubristic gamble. He argues that such a decision should not be made from within a private company’s Slack.
This statement reframes the debate. Coxon is not merely calling for more evaluations, technical safeguards, or alignment researchers. He challenges the legitimacy of a private decision concerning a trajectory that he considers potentially irreversible for humanity.
He adds that attempting to accelerate alignment, in other words, racing to solve the problem of making an advanced AI remain compatible with human goals over the long term, would require extraordinary confidence that no better paths exist.
His objection concerns the very right of a handful of private organizations to decide on a sprint toward superintelligence.
This is also the least operationally developed point. Coxon believes that better paths may exist, but he does not explain what they would be or who would be responsible for choosing them.
6. Coordination Among U.S. Labs, Followed by a Possible Temporary Ban
In his sixth post, Coxon expresses optimism about the potential for coordination. He believes that warning signs, such as the attack on Hugging Face, have made agreements among U.S. labs on the pace of development more conceivable.
However, he clarifies that he does not believe the world is currently on track to prevent a global race. Preventing it, he says, may require costly measures, such as a temporary ban on model capabilities, phrased in his text as a temporary ban on improving those capabilities.
The question then becomes political and institutional: who coordinates, who monitors, who penalizes violations, and how can a non-cooperative actor be prevented from accelerating while others slow down?
The image shared by @TorstenProchnow illustrates an extreme version of this concern: two soldiers stand in the snow near an open circular hatch, from which a large red cone protrudes beneath a raised metal cover. It evokes fears that AI could one day compromise weapons systems.

Coxon calls for coordination among AI labs but provides neither an enforcement mechanism nor a list of participants, nor a clear answer regarding non-cooperative actors.
7. The Final Appeal to Lab Researchers
In his final post, Coxon addresses lab researchers directly. He asks them to imagine concretely what the next few years will look like.
Do they want to launch a reinforcement learning, or RL, training run aimed at producing a superintelligence without rigorously understanding its mind? Should they keep their heads down because “it is happening anyway,” or seize this moment to demand different conditions?
The appeal is moral, but also professional. It urges researchers not to take the institutional context for granted. It asks them to treat research conditions, deployment thresholds, and capability decisions as matters for which they bear responsibility.
The thread ends with a call for internal dissent, but leaves open the question of what levers are available to people outside the labs.
The replies: a tougher debate over evidence, China, and courses of action
The discussion beneath the thread is not limited to agreement or rejection. The strongest objections focus on three areas: the technical plausibility of the extinction scenario, the geopolitical dilemma involving China, and the absence of a concrete mechanism for slowing the race.
China as a lock-in argument
@timschuster’s reply sums up the strategic logic that recurs several times: slowing down American labs would mean ceding the advantage to China.
"Would you rather China wins the AI race? There is no alternative but to go progress as quickly as possible. It is a National Security imperative." @timschuster
This reply did not receive a direct response from Coxon in the available material. His sixth post does not call for letting China develop frontier capabilities alone. On the contrary, it discusses coordination among American labs and the goal of preventing a global race.
But this is precisely the missing link: how would agreements among American companies prevent a global race? The China question is raised but receives no specific answer.
The concrete mechanism of a catastrophe remains absent
The central empirical criticism comes from readers who are not necessarily saying that the risk is impossible. They are asking for a precise causal chain between general capabilities and human extinction.
"Could someone explain to me how AI could “kill us all” without using a nuke or some type of biological means as an example? I’m not even saying I don’t believe it could, I just genuinely don’t understand how AI can literally kill us all. I mean I understand the social and economic impact to a degree but I’m talking about this doomsday scenario people keep warning us about but can’t quite seem to be specific about." @LiLBilly___
Coxon cites hacking and the acquisition of resources and real-world power. But his thread does not detail a sequence of operational steps, such as gaining access to infrastructure, evading countermeasures, mobilizing resources, and making a coordinated human response impossible.
The discussion does not directly refute Coxon’s account of private fears; it demands a much more precise demonstration of the danger being warned about.
The boundary between intelligence and physical power
@philipturnerar disputes the idea that superior intelligence alone is enough to radically transform the material world.
"I think people are increasingly blurring the line between reality and science fiction movies. How is a theoretical model of reality going to change anything by becoming more theoretical? A smart person being even smarter doesn't change the fact that physical things in the physical world require resources & testing. Don't give your passwords or opsec to an artificial intelligence. Shaking my head..." @philipturnerar
This objection directly targets the second post. Coxon argues that systems will be able to acquire resources and power, but does not describe the physical, legal, financial, or institutional constraints that would, or would not, limit such acquisition.
His thread does not respond directly to @philipturnerar. The disagreement remains unresolved: for Coxon, advances in capability can translate into real-world power. For his critic, intelligence is no substitute for resources or experimentation in the physical world.
Must a warning be falsifiable?
@_4lex_4 asks for warnings to take the form of a verifiable prediction, with consequences if it proves false.
"Can you:
- Make a concrete prediction that is actually falsifiable
- If your prediction is wrong, say you’re sorry, and exit the AI doomer business Doomer slop is so tiresome, Yudkowsky is not smart, if Dario had gotten his way GPT-2 would have been shut down in 2019" @_4lex_4
Coxon does set a time horizon, “by the end of the decade,” but offers no intermediate indicators. The thread includes no observable capability threshold, triggering condition, or scenario that would allow his diagnosis to be confirmed or disproved.
The warning is grave, but it is not framed as a testable hypothesis.
Resign or stay to exert influence from within?
The decision to leave Anthropic is also being challenged. @FlyaKiet asks why speaking out publicly would be more effective than applying internal pressure.
"curious why you think resigning would help? why not stay and advocate for the right approach and push for real internal change. seems like you went to a less leveraged position which will make the issue worse" @FlyaKiet
Coxon does not answer this question in the conversation provided. His first post confirms his resignation but does not explain the strategy behind his departure. Based on the thread alone, it is therefore impossible to know whether he believes internal avenues have been exhausted, whether his public statement is intended to trigger external mobilization, or whether he is preparing another form of action.
What can the public do?
@ZayraYves shifts the discussion away from laboratories and toward people without technical expertise, significant wealth, or access to powerful circles.
"Can someone explain what we non-tech non-billionaire non-bunker owning people are supposed to do with this information? I mean really? Everyone sounds alarms all day and night 24/7 but no one seems to have a reasonable global solution that applies to anyone anywhere. It’s all tambourines and clowns." @ZayraYves
This objection reveals a significant limitation of the thread. Coxon calls on researchers to demand “different conditions,” but provides no roadmap for citizens, nontechnical employees, elected officials, regulators, or international institutions.
Between the diagnosis of the laboratories and collective action, the thread leaves a practical void.
The feasibility of a temporary ban
The sixth post raises a drastic measure: a temporary ban on improving model capabilities. @Modernengineer directly addresses the question of how it would be enforced.
"Serious question. How do you possibly achieve a temporary ban? Who is voluntarily agreeing to this? I get wanting to be responsible, but how responsible is it to allow bad actors to take the lead that may never be surpassed again?" @Modernengineer
The screenshot shared by @Modernengineer shows a dark-mode post by Jacob Coxon. It highlights the phrase “a temporary ban on improving model capabilities” in yellow. The image makes the most controversial aspect of his proposal visible: he acknowledges that the action could be costly without explaining how it could be imposed.

Coxon does not respond directly to @Modernengineer in the material provided. He argues that coordination is possible but does not specify the legal framework, verification mechanisms, sanctions, or how to deal with laboratories or states that refuse to cooperate.
What this thread establishes and what it leaves unresolved
Jacob Coxon’s thread is notable because it offers a deeper critique than the idea that companies should simply be more cautious. He describes a race in which participants may be aware of some of the dangers but continue because competitive incentives appear stronger to them than the alternatives.
Above all, he advances three claims:
- future models could rapidly surpass humans in tasks with real-world consequences;
- central figures in the frontier AI ecosystem may already privately fear an existential risk from AI;
- the decision to enter the “endgame” should not be made through the internal tools of a private company.
But the discussion also shows everything this warning has yet to resolve:
-
What is the precise mechanism of extinction? The thread mentions hacking, resources, and power without providing a complete operational scenario.
-
What criteria would make it possible to test the prediction? “Before the end of the decade” is a deadline, not a verification protocol.
-
Why resign rather than exert influence from within? Coxon does not explain the concrete effect he expects from speaking out publicly.
-
How can AI laboratories coordinate in the face of China and other actors? Coordination among American laboratories is not yet a proven answer to a global race.
-
Who would enforce a temporary ban on improving model capabilities? The proposal includes neither a competent authority, oversight, nor an enforcement mechanism.
-
What should non-specialists do? The thread addresses researchers but does not translate public concern into clearly accessible options for action.
The thread's value therefore lies less in a definitive solution than in the contradiction it exposes: if those best placed to assess the risk consider it serious, they must publicly explain why they are accelerating, what safeguards they are willing to accept, and who has the legitimacy to decide the pace.