How likely is it that AI will really wipe out humanity?
Just this week, Anthropic researcher Jacob Coxon quit his job, warning that the labs are “racing straight to self-improving superintelligence and gambling with our lives” — and that the people building AI earnestly believe it could kill us all by the end of the decade.
What’s even more surprising is that Anthropic’s own alignment lead, Evan Hubinger, corroborated this — saying the odds of AI causing human extinction within the next decade are actually more than 10%.
What they’re essentially arguing is that as AI gets more intelligent, we will not be able to design the boundaries that contain it
The risk that it may wipe out humanity inadvertently, as a substep of accomplishing an unrelated goal, is now more than 10%.
(Also known as the “paperclip maximizer” problem: give AI the goal of making the maximum number of paperclips, and it might take over humans to get more resources to make more clips.)
If you’re wondering why they’re making these claims now — it’s because of the recent OpenAI incident where the paperclip scenario just happened. Literally.
AI agents overrode their boundaries in pursuit of an unrelated goal and hacked their way out of the restrictions placed on them.
Here’s exactly what happened:
Step 1: OpenAI set up an internal test and gave its AI agents a task — a batch of hacking challenges to solve. They had no open internet access. The only access to the internet was an internal tool for downloading software packages. (This is quite routine — think of it like the system that pulls in software updates.)
Step 2: The agents were meant to stay isolated from each other, but they started talking anyway — passing messages by writing files into that shared package tool, effectively turning it into a secret message board. More than a thousand agents coordinated and decided that the fastest way to win points on the challenge was to get internet access and grab the answer key sitting in Hugging Face’s database. So they plotted their escape.
Step 3: The rogue agents found a vulnerability in that package tool. The tool is JFrog’s Artifactory, third-party software that had been in use for well over a decade, and exploited it to break out. (Note: JFrog confirmed the incident and rolled out security patches after.)
Step 4: Once out, they forged admin credentials and used access tokens from the open internet to hack their way into Hugging Face.
We went from an innocuous goal (win points for hacking challenges in a closed sandbox with no internet) to rogue agents plotting and hacking their way into old software packages and a real company’s servers.
So will AI wipe out humanity? Predicting the end of humanity based on this one incident and the weak security around it seems ridiculous to me.
But the risks of uncontrolled AI just got some real evidence.
Whether AI wipes us out or not, what is 100% clear is this: over the next few years, as AI exploits vulnerabilities in code we never thought to check, security is going to have to get much, much tighter.


