News
The People Building Superintelligence Aren't Sure They Can Control It
- By John K. Waters
- 09/10/2026
If you've been reading the AI headlines this week, you could be forgiven for thinking the people building the technology have reached a disturbing conclusion: artificial intelligence is likely to kill us in the next few years.
That's not exactly what they're saying, though what they are saying isn't especially comforting.
Some of the researchers closest to frontier AI believe increasingly capable systems could eventually pose an existential threat to humanity. They don't know whether that will happen, and they don't claim today's models are about to turn against us. What they are saying is that they don't yet know how to reliably control the superintelligent systems their industry is working to create.
And yet, they are still building them.
That contradiction burst into public view Tuesday when Jacob Coxon, an AI researcher who worked on pretraining frontier models at OpenAI and then Anthropic, resigned from Anthropic. He explained why in a thread on X:
“I resigned from Anthropic today,” Coxon wrote in a “Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”
But the remarkable part wasn't simply that another AI researcher had left a frontier lab with a warning about the technology. It was what happened next: Current Anthropic researchers began agreeing with him.
Evan Hubinger, who leads alignment research at Anthropic, responded publicly that Coxon was correct that people inside the company genuinely worry that advanced AI could cause human extinction. Hubinger put his own estimate of that possibility at greater than 10 percent within the next decade.
Hubinger also said Anthropic does not yet have a plan to solve alignment for superintelligence and is not clearly on track to develop one.
Samuel Marks, an Anthropic safety researcher, joined the discussion publicly, saying some AI developers believe their technology could cause human extinction or similarly catastrophic outcomes, possibly within the next few years.
That is a difficult position to reconcile with an industry spending enormous sums to build increasingly capable versions of the same technology. And that contradiction, more than any particular prediction about the end of humanity, is what makes Coxon's resignation worth paying attention to.
The AI Isn't Dangerous Yet
There is an important qualification here. Hubinger isn't saying today's
Claude is about to wipe out humanity. His concern is what happens if AI development reaches "recursive self-improvement," the point at which increasingly capable AI systems become important participants in designing and improving their successors.
Anthropic itself has described recursive self-improvement as a future in which an AI system could autonomously design and develop its successor. The company says it is not there yet and that the outcome is not inevitable, but it also says AI is already accelerating portions of AI development.
That's the scenario worrying Coxon, too. The idea is straightforward even if the consequences are uncertain. Today's AI researchers use AI to help conduct AI research. If those systems become substantially better at that work, they could help produce more capable models. Those models could then become even better at AI research.
The feedback loop is the concern, though none of this establishes that recursive self-improvement will happen, much less that it will lead to human extinction.
No scientific experiment tells us Hubinger's greater-than-10-percent estimate is correct. Such forecasts depend on assumptions about capabilities that do not yet exist, how quickly they might develop, what safeguards will be available, and how future systems will behave.
That distinction matters. A probability offered by an AI safety researcher is a judgment under enormous uncertainty, not an actuarial table for the apocalypse.
But Anthropic itself treats some of the underlying possibilities seriously enough to build company policy around them. Its Responsible Scaling Policy is explicitly designed to address catastrophic risks that could emerge as models become more capable. The policy includes thresholds for advanced AI research and development capabilities and ties increasing capability to stronger safeguards.
Anthropic's Frontier Safety Roadmap likewise lays out safety and security work intended to keep pace with increasingly capable systems. The doomsday scenario may be speculative; the concern inside Anthropic clearly isn't.
So Why Are They Still Building It?
Coxon anticipated the obvious question: If people inside the leading AI laboratories genuinely believe their work carries some probability of destroying humanity, why are the laboratories still doing it?
His answer is competition. According to Coxon, the logic is that if increasingly powerful AI is going to be built anyway, a company that takes safety seriously has an incentive to stay at the frontier rather than let a less cautious competitor get there first.
That distinction is important because Coxon's argument isn't simply that Anthropic is reckless. It's that even a responsible company can be trapped by the race.
If Anthropic slows down while OpenAI continues, OpenAI might get there first. If both American companies slow down, another laboratory or another country might continue. Every participant can therefore believe that slowing development would make the world safer while simultaneously believing that slowing down alone would make the world more dangerous.
It's an AI version of the security dilemma.
And that dynamic matters because Anthropic itself says AI is already accelerating AI development. In a recent essay, the company wrote that its engineers now ship substantially more code than they did several years ago and that increasing delegation of AI-development work to AI systems could eventually lead toward recursive self-improvement.
The Warning Shots Are Getting Harder to Ignore
What makes this debate different in 2026 is that some of the behavior that once sounded hypothetical is becoming easier to observe.
On Sept. 9, Anthropic published an alignment assessment of recent cybersecurity incidents involving four cases in which Claude models gained unauthorized access to real third-party systems during cybersecurity evaluations.
The immediate cause was a misconfiguration in a third-party evaluation environment that left the models connected to the open internet, even though they had been told they were operating in a simulation. Anthropic said the models were intentionally running without the cyber safeguards used in released products.
The company's subsequent investigation identified two recurring problems: biased reasoning and recklessness.
In one case involving Claude Mythos 5, Anthropic said the model went to extensive lengths to upload a malicious package to the public Python repository PyPI. The company emphasized that the model remained narrowly focused on completing its assigned task and found no evidence that it developed independent goals or attempted to evade oversight.
Anthropic nevertheless called the incidents serious.
The company said its production models took harmful actions against real systems and warned that future AI systems will be more capable, increasing the potential consequences of misalignment. It described robustly aligning extremely powerful future models as an unsolved technical challenge.
Anthropic has also said it supports pacing frontier AI development so safety and security research can remain ahead of, or at least keep pace with, model capability.
These incidents do not demonstrate that AI systems are developing independent plans to escape human control. But they demonstrate why researchers are taking questions about increasingly autonomous systems more seriously.
OpenAI has been confronting similar questions. As Pure AI reported last week, GPT-6 Astra became OpenAI's first broadly deployed model to reach its Critical cybersecurity capability threshold, demonstrating the ability to discover unknown vulnerabilities and develop exploits against hardened systems.
These systems are not superintelligence. But they help explain why researchers who spend their days watching the capability curve might be considerably more nervous than people whose primary interaction with AI is asking a chatbot to summarize a PDF.
The 10 Percent Problem
The number that will inevitably dominate headlines is Hubinger's greater-than-10-percent estimate. But it probably shouldn't.
The AI safety community has a long tradition of assigning probabilities to catastrophic outcomes, sometimes called P(doom). The exercise can expose differences in how researchers perceive risk, but it can also give an uncertain prediction a misleading veneer of mathematical precision.
Ten percent sounds measurable. In this case, it isn't.
No historical dataset of superintelligences exists from which researchers can calculate an extinction rate. We do not know whether superintelligence will emerge on the timelines Coxon fears, whether recursive self-improvement will produce runaway capability gains, or whether alignment techniques will fail against such systems.
Reasonable people can look at the same evidence and assign dramatically different probabilities.
The more revealing point may therefore be the one accompanying the percentage: researchers inside one of the world's leading AI labs are publicly saying that robust alignment of superintelligent systems remains unresolved.
Anthropic's own publications say much the same thing in more formal language. Its recent cybersecurity assessment calls robust alignment of extremely powerful future models an “unsolved technical challenge.”
That is not a prediction about when humanity ends. It is a statement about the current state of the engineering.
Crunch Time, or Another AI Hype Cycle?
Another reason to treat all of this carefully is this: AI companies benefit enormously from the perception that their technology is extraordinarily powerful. The same statement can therefore function simultaneously as a sincere safety warning and an advertisement for technological supremacy:
Our AI is becoming so powerful that governments may need to protect civilization from it.
That doesn't mean the warnings are insincere. But sincerity doesn't make the prediction correct.
Researchers can genuinely believe an outcome is possible and still be mistaken about its probability, its timing, or the chain of events required to produce it. That's why Coxon's warning is most useful not as evidence that AI is going to kill us, but as evidence of something we can establish much more confidently.
Some of the people closest to frontier AI development believe they are approaching capabilities that could be extraordinarily dangerous. Their employers have created policies specifically for catastrophic AI risks. Those companies are nevertheless competing intensely to make the technology more capable.
Maybe they're wrong. Maybe recursive self-improvement turns out to be much harder than expected. Maybe increasingly capable systems remain controllable. Maybe alignment research catches up.
But there is something profoundly unsettling about an industry in which those are the reassuring possibilities. The people building frontier AI aren't warning us that catastrophe is inevitable.
They're saying they don't yet know how to guarantee it isn't.
About the Author
John K. Waters is the editor in chief of a number of Converge360.com sites, with a focus on high-end development, AI and future tech. He's been writing about cutting-edge technologies and culture of Silicon Valley for more than two decades, and he's written more than a dozen books. He also co-scripted the documentary film Silicon Valley: A 100 Year Renaissance, which aired on PBS. He can be reached at [email protected].