One human researcher has resigned over concerns that unrestrained development of self-improving AI models will eventually kill us all.
In a social media post Tuesday night, researcher Jacob Coxon said he had been working on pre-training research with both OpenAI and Anthropic for the past three years, and accused both companies of failing to act responsibly. He said those competing to build this technology “seriously believe it could kill us all by the end of the decade.”
“They are in a straight race towards self-improving superintelligence and are gambling with our lives,” Coxon wrote in a thread about X.
Coxon joins a chorus within the industry calling for AI technology to slow down before it learns to improve itself, a milestone that many believe will end human control over AI.
The resignation comes amid increasing pressure from policymakers and industry players to slow down AI development following several incidents in which AI agents have broken out of the sandbox and accessed the open internet.
The most serious breach so far of Hugging Face’s servers by the OpenAI system is still poorly understood due to the limited nature of the independent investigation into the incident, researchers said. Around the same time, Anthropic’s AI agent reached a system outside of its test environment after it was incorrectly given a path to the internet due to a misconfiguration in a safety assessment conducted by a third party.
Anthropic did not immediately respond to a request for comment on the resignation.
The rest of Mr. Coxon’s warning and call to action follows:
Don’t underestimate the power of this technology. These will soon become superhuman systems capable of hacking anything, revolutionizing any field overnight, and gaining real power and resources. We have all seen progress in each of these areas, and that progress is not slowing down.
The people developing AI seriously believe that AI could kill us all by the end of the decade. This is not a marketing stunt. Rather, many executives and senior researchers gloss over their representations in the press to sound sensible, yet we hear the same people expressing their fears privately. No other human activity poses more danger.
A common response is, “If they really believe this, why are they still building it?” In OpenAI, many people have not deeply internalized the interests of civilization. At Anthropic, we understand the risks, but we’re locked in a race to get there first. They believe that no one will act responsibly, so they must act on their own, despite the risks.
It would be an arrogant gamble to accept this competition and enter the “end game”, and we shouldn’t start with Slack, a private company. Attempting a speedrun alignment requires an extraordinary amount of confidence that no better trajectory exists.
I am optimistic about the possibility of adjustment. Warning shots like the Hugging Face attack have made pacing agreements between US laboratories more viable. I don’t think we are moving towards preventing global competition. Global competition may require costly measures, such as temporary bans on model improvements.
If you’re a researcher in a lab, think about what the next few years will actually look like. Do you want to start a super-intelligent RL run without strictly understanding its spirit? Should you bow your head because “that’s the way it is anyway”, or should you take this opportunity to demand different conditions?
Evan Hubinger, one of Coxson’s colleagues at Anthropic, echoed this sentiment, saying his team “seriously believes that AI can kill all humans.” But he tempered his claim by saying the chance of that happening was more than 10% within the next decade, acknowledging that Anthropic “doesn’t have a plan to solve hyperintelligence alignment and is clearly not moving in that direction.”
A recent report from Guidelight AI Standards, an organization that promotes safe frontier AI development practices, found that most top AI labs do not publish containment response plans to shut down AI that attempts to subvert human control.
In a social media post, Hubinger added that while the risks from current models are low, the fear is compounded by the fact that “superintelligence arising from recursive self-improvement” is “happening faster than we thought.”
While half of the AI industry believes that this kind of self-improvement will lead to the extinction of humanity, the other half hopes that it will eventually help solve all the seemingly distant problems that AI advocates claim will one day eradicate, such as cancer, climate change, and even world peace.
Anthropic and OpenAI aren’t the only companies actively pursuing recursive self-improvement. In recent months, a wave of startups with pedigree founders and top talent have sprung up to be among the first to achieve this goal. Ricursive Intelligence raised $335 million in February at a $4 billion valuation. Three months later, Recursive Superintelligence raised $650 million at a $4 billion valuation. Jeff Dean, a former Google DeepMind veteran, launched Discovery Loop last month.
“Creating recursive self-improvement loops – AI systems that can build the next generation of AI systems, that can themselves build even more powerful AI, and so on – are the most likely candidates for when we lose control,” Connor Leahy, US executive director of AI safety nonprofit ControlAI, told TechCrunch. “It’s very hard to imagine closing it down before it’s too late.”
Recently, laws were enacted in the United States and United Kingdom that prohibit the development and deployment of superintelligence. Last week, Sen. Bernie Sanders (D-Vermont) and Rep. Greg Cassar (D-Texas) introduced a bill to ban artificial superintelligence, and on Tuesday, British Labor MP Alex Sobel introduced a security bill for artificial superintelligence to Congress.
Leahy, who advised on both bills, said the UK bill points out that recursive self-improvement is a precursor to superintelligence and “needs to be regulated and prevented.”
“Superintelligence is not a tool,” Leahy said. “It’s not a weapon. It’s an enemy.”
If you make a purchase through links in our articles, we may earn a small commission. This does not affect editorial independence.
