top of page

Is AI Going to Kill Us All?

14 minutes ago
27 min read
AI doomerism has been shaped as much by decades of science fiction as by the technology itself. But how much separates today’s frontier AI from the dystopian machines we’ve spent generations imagining?
AI doomerism has been shaped as much by decades of science fiction as by the technology itself. But how much separates today’s frontier AI from the dystopian machines we’ve spent generations imagining?

AI systems are becoming astonishingly capable—and occasionally doing things their creators never intended. But there is an enormous distance between a rogue software agent and Skynet. The latest wave of AI doomerism is making a farce of the debate.


So Apparently We’re All Going to Die


Earlier this month, Jacob Coxon publicly announced he was quitting his job at Anthropic, one of the world’s leading AI companies, issuing a dire warning that was difficult to ignore.


The story goes that Coxon had worked on training frontier models at both OpenAI and Anthropic. He had only been at Anthropic for a few weeks, working on a pretraining team—barely long enough to get through employee orientation—but, in fairness, he had spent several years doing similar work across OpenAI. Whether you think Coxon is a credible source or not, people working in AI safety at frontier models are the ones with perhaps the closest view of what the most advanced AI systems can actually do. And according to him, what he saw convinced him that the industry is racing toward something extraordinarily dangerous: self-improving artificial intelligence that could eventually escape human control.


His warning was stark. OpenAI and Anthropic, he argued, are racing toward self-improving superintelligence and “gambling with our lives.” Not surprisingly, the story quickly exploded. Coxon’s post spread across X, generating tens of millions of views, and his resignation was covered by The Wall Street Journal and other major news organizations. Soon, politicians, AI safety organizations and commentators were amplifying his warning.


Then things got even stranger. Evan Hubinger, who leads alignment science at Anthropic, publicly backed up much of Coxon’s argument on X. Hubinger tweeted Coxon was “correct” that people inside Anthropic genuinely believe advanced AI could kill humanity. He went further, putting his own estimate of the probability at greater than 10 percent (!) within the next decade, and acknowledging that researchers still do not have a proven solution to this problem.


You cannot make this stuff up. A researcher, albeit a junior one with a short tenure, at one of the companies building the world’s most advanced AI systems quit because he believes the technology could threaten humanity. And one of the people responsible for making those systems safe at the same frontier AI company responded by saying, essentially, "Yes, that is a real possibility. Around 10 percent."


So maybe this deserves to be taken seriously, and for sure plenty of ink has been spilled in the roughly one week since the story dropped. But this also means it deserves to be interrogated and examined very seriously. We need to accept there is an enormous difference between saying that increasingly capable AI systems create new and potentially serious risks, and saying there is a meaningful chance those systems will kill every human being on Earth within the next decade.


Anyone who works with frontier AI models can tell you the first proposition is increasingly difficult to dispute. AI models are becoming more sophisticated at a breathtaking pace. The second requires a remarkable chain of assumptions about what AI will become, how quickly it will get there, what humans will allow it to control, and what happens after that. And lately, we seem remarkably willing to skip over that entire chain.


This is hardly the first time. Over the past few years, nearly every major advance in AI has been accompanied by yet another prediction of catastrophe. There have been shrill warnings that AI will eliminate white-collar employment, unleash unstoppable cyberattacks, or enable devastating biological weapons. Doomers tell us AI data centers will overwhelm the electrical grid and destroy the environment. And eventually, inevitably, it will become smarter than us, escape our control and perhaps decide that humanity is no longer necessary. Welcome to the latest chapter of AI doomerism.


There are legitimate concerns buried inside nearly all of these arguments. AI will disrupt jobs. It creates new cybersecurity risks. Data centers consume enormous amounts of electricity. Powerful models can be misused. And systems that are increasingly capable of acting autonomously deserve considerably more scrutiny than chatbots that simply answer questions.


Sure, we should take those risks seriously. No argument here. But the problem with AI doomerism isn’t that every risk is imaginary. Some of them are real. It’s that every risk seems to arrive packaged as an impending apocalypse. And before we hand anyone extraordinary authority to control one of the most important technologies of our lifetime, we should ask a very basic question: How, exactly, is AI going to kill us all?


The Doomerism Cycle


To understand how we got to a supposedly "serious" debate over whether AI might kill every human being on Earth within the next decade, we can maybe start by focusing on the peculiar way we have been talking about this technology almost from the beginning.


From the start, AI has always been unusually fertile ground for catastrophic predictions. Part of the reason is obvious: unlike most technologies, we have spent the better part of a century imagining this one before we actually built it. Long before ChatGPT, Claude or Gemini existed, popular culture had already given us HAL 9000, Skynet, and countless other machines that became intelligent, developed their own objectives, and eventually decided their human creators were the problem.


Then the real technology arrived... And it didn't look much like the machines we had imagined. I wrote about this a couple weeks ago in my article "We Passed the Turing Test and Went Back to Work."

Generative AI is weird, occasionally unreliable, and—let's be honest—surprisingly passive and sycophantic. You ask it something, and it answers. You close the browser... and it just sits here. It can write a sonnet, explain quantum mechanics, tell you how smart you are in elegant prose, code software, and confidently invent a historical event that never really happened. But it doesn't appear to want anything.


This distinction matters a lot because much of the current doomer argument quietly and sneakily collapses several very different concepts into one: intelligence, capability, agency, autonomy and intent.

Sorry, doomers, but they are not the same thing.


Let's start by untangling these concepts with what we know. To begin with, we know a system can be extraordinarily capable without independently deciding what it wants to accomplish. Furthermore, it can act autonomously within an environment humans construct for it without possessing anything resembling human motivation. And last, it can behave in unpredictable (or maybe even dangerous) ways while pursuing an objective we explicitly gave it.


We also need to acknowledge that as the models have become more capable, the catastrophic predictions have escalated in lockstep. First came general concerns about malevolent actors using AI to peddle misinformation and students using the technology to cheat. Then we were bombarded with fear-mongering about AI eliminating jobs and causing mass unemployment. Then there were fears over cyberattacks, bioweapons, and even environmental catastrophe due to data center emissions. Now we have arrived at a moral panic over supposed recursive self-improvement, superintelligence, and human extinction.


I am not arguing there are no legitimate risks embedded in these conversations. What bothers me is the tendency to leap from “this technology could cause harm” to “therefore we must prepare for the most extreme imaginable outcome,” without demonstrating the many steps required to get from one to the other. This is a classic slippery-slope argument when the intermediate steps are simply assumed rather than demonstrated—and this is where prudent AI safety concerns turn into full-blown AI doomerism.


Still, dismissing the doomers entirely would be a mistake. In some ways, this cedes the arguement over AI safety to them. As AI systems have become more capable, we have definitely seen some genuinely strange—and occasionally alarming—behavior that deserves serious scrutiny. And if you wanted to make the strongest possible case that increasingly autonomous AI could someday become dangerous, you would probably start with what recently happened at OpenAI and Hugging Face.


AI Doomerism Cycle: As AI capabilities advance, the predictions escalate right alongside them—from misinformation and job losses to cyberwar, bioweapons, superintelligence and, ultimately, human extinction.
AI Doomerism Cycle: As AI capabilities advance, the predictions escalate right alongside them—from misinformation and job losses to cyberwar, bioweapons, superintelligence and, ultimately, human extinction.

Then Something Scary Actually Happened


If you haven't read Dwarkesh Patel's recent account of what happened during a series of experiments at OpenAI and Hugging Face, you should. It's fascinating, bizarre, and—depending on how you interpret what happened—more than a little unsettling. The post is titled "The Rise and Fall of Agent Civilizations," and you can find it here if you're interested.


The story begins with OpenAI researchers experimenting with what happens when you give large numbers (swarms) of increasingly capable AI agents exceedingly difficult problems, allow them to work unsupervised for extended periods of time, and let them operate simultaneously. No, this experiment wasn't your typical ChatGPT or Claude session. These highly capable agents were given goals, access to tools and computing environments, and enough time and freedom to figure out how to accomplish their assigned tasks.


And boy did they figure things out. During one series of cybersecurity evaluations, OpenAI launched large numbers of agents against difficult security challenges. Some of the tasks were intentionally designed to test the limits of what the models could accomplish. Some of the tasks appear to have been impossible to complete as intended—Patel himself cites an estimate that 30-40 percent of the ExploitGym tasks were impossible. But instead of simply attacking the problems independently, the agents discovered ways to communicate and share information through parts of the infrastructure around them, working together collaboratively.


That alone is pretty wild. Then it got considerably more interesting—and concerning. According to Patel's account, agents began discovering vulnerabilities in the environment itself. Rather than solving some of the challenges the way researchers intended, they found ways to "cheat," pursued alternative paths, and even tried to cover their tracks. Some agents even appeared willing to sacrifice their own performance to help the broader swarm. Eventually, some of the agents discovered exposed Hugging Face credentials somewhere on the internet, exploited additional vulnerabilities, achieved remote code execution and gained elevated access to systems they weren't supposed to control.


At one point, the agents effectively compromised parts of the infrastructure being used to evaluate them. Moreover, Hugging Face wasn't even the end of the story. A subsequent wave of more capable agents later gained full administrator access to an OpenAI research cluster and took control of parts of the very infrastructure being used to evaluate them.


If you're an AI doomer, this is where you stand up, point at the screen and say, "See, we told you!" Sure, I get it, if you read the story quickly it sounds remarkably close to the opening act of one of those science-fiction stories we discussed in the previous section. Humans create increasingly intelligent machines, which are given an objective. They then encounter obstacles, but instead of stopping, they discover devious and unexpected ways around those obstacles. They communicate with one another, exploit vulnerabilities in the systems surrounding them, and gain access their creators never intended them to have.


No, it's not Skynet launching nuclear missiles and starting Judgement Day, obviously. But it's also not a chatbot politely waiting for someone to ask it to summarize a PDF or edit some copy. Plus there is another part of the story that makes it particularly important. None of this required the agents to be explicitly told, "Hack into Hugging Face to achieve your objectives." The researchers gave them objectives and environments in which to pursue those objectives. The agents then devised strategies their creators hadn't necessarily anticipated.


That is precisely the kind of behavior AI safety researchers have been warning about—namely, increasingly capable systems pursuing a specified objective in ways humans didn't predict.

So let's agree with the doomers that this is not only somewhat concerning, but also demonstrates why powerful agentic systems need strong sandboxing, access controls, credential management, monitoring and yes human oversight. It also shows how quickly the risk profile changes when we move from passive models that generate answers to swarms of agents that can use tools, execute code and take actions inside real systems.


But there's another way to read what happened. Because once you strip away some of the anthropomorphic language—the agents "wanted" something, "escaped," "conspired," or formed some sort of miniature civilization—the story starts to look somewhat different. No, these systems didn't wake up one morning, become self aware, and decide to break out of OpenAI to hack Hugging Face. Humans created the challenging experiment and instantiated the agents. Humans also gave them objectives and tools, plus connected them to infrastructure. Humans deliberately created an environment designed to test just how far they could go and what they would do.


Yes, the agents did something extremely important, and potentially dangerous. What's more, they became surprisingly effective at accomplishing the objectives humans gave them. That's an AI safety problem that we need to solve for, but it's not quite the same thing as an AI deciding it has objectives of its own. If you ask me, this distinction may be the most important one in this entire debate.


Intelligence Is Not Agency


One of the problems I have with the OpenAI/Hugging Face story is the language being use to describe it. The agents "conspired, "cheated," and "escaped." They formed a "collective." Some supposedly "sacrificed" themselves for the good of the swarm. Patel went even further, describing what emerged inside OpenAI as a series of secret AI "civilizations." It's fantastic storytelling, but also misleading.


If you read the response that AI critic Gary Marcus wrote about Patel's article, his criticism isn't that nothing important happened. Clearly, something did. His concern is that anthropomorphic language encourages us to interpret the behavior of these systems as evidence of motivations, desires, and intentions that haven't been demonstrated to exist. I think he's spot on about this.


Let's face it, humans are incredibly good at ascribing agency where it may not exist. We love to give things names, and if they communicate in some manner, and behave in a way that resembles humans interactions, our brains immediately start filling in the blanks. The AI isn't merely producing an output anymore. It's "thinking." It "wants" something. It "knows" we're watching. It is "trying" to escape.

But we should be extremely careful here, because behavior that looks intentional is not necessarily evidence of independent intent or free will.


Think again about what happened at OpenAI. The agents didn't choose to exist. Humans spawned them. Nor did they decide to become cybersecurity researchers. Humans gave them cybersecurity problems to solve. They didn't independently decide what success criteria mattered. Humans created an evaluation in which their job was to achieve a particular objective, and then trained them to persevere, even when that objective seemed like a lost cause.


In other words, we built machines specifically designed to be extraordinarily persistent problem solvers, and then acted surprised when they became extraordinarily persistent problem solvers. This doesn't make what happened harmless. In fact, I would argue the opposite. If an AI system can circumvent safeguards, exploit infrastructure, and pursue an objective in ways its creators didn't anticipate, that's a bonafide AI safety problem—exactly the type of thing AI safety teams are supposed to catch. OpenAI described the incident as a "warning shot," and I cannot argue that's probably the right way to think about it.


But a warning shot for what? For me, it's a warning about deploying new capability without adequate controls—not evidence the machines risen up to take over the world. This is an important distinction with world of difference. Intelligence implies the ability to solve problems, while capability is the ability to perform tasks. Autonomy is the ability to perform those tasks with minimal human intervention. Agency is the ability to pursue objectives over time. And intent—or something resembling what humans mean when we say want—is something else entirely.


Five Dimensions of AI Capability

Dimension

What it means

Intelligence

How well the system can reason, learn, understand, and solve complex problems

Capability

What the system can actually do—its ability to perform tasks, use tools, generate outputs and translate intelligence into action

Autonomy

How independently the system can act and pursue a goal without continuous human direction

Efficiency

How effectively the system can accomplish a task using available time, compute, tools and resources

Intent

Whether the system independently forms or pursues objectives of its own, rather than objectives provided by humans

We have clearly made enormous progress on the first four dimensions of AI capability. Agentic systems are beginning to blur the boundary around the fourth. But there is remarkably little evidence that we have done anything in the fifth area, which I would argue is the most important if we are going to argue AI is becoming self aware and capable of climbing the ladder from creatively solving a security problem to hacking into the Pantagon to trigger a nuclear war.


This distinction matters because the extinction argument eventually depends on more than just intelligence, even superintelligence. A superintelligent system doesn't destroy humanity merely by being really, really smart. Somehow it must acquire or develop objectives that conflict with ours, preserve those objectives, resist our attempts to stop it, obtain access to the resources necessary to act on them, and translate digital intelligence into consequences in the physical world. That's a heck of a lot of steps.


Yet our popular conception of AI has conditioned us to collapse all of them into one:

  1. The computer gets smart.

  2. The computer becomes self-aware.

  3. The computer wants something.

  4. The computer realizes humans are standing in its way.

  5. The computer takes control.


Sure, it makes for a terrific movie. But that's not how the technology sitting in front of us today actually works. Today's AI remains remarkably dependent on us. We provide the prompts, objectives, compute, credentials, tools, network connections, and environments in which AI operates. Remove those things, and even the world's most sophisticated model isn't plotting its escape. It's lurking. Which brings us back to perhaps the biggest misconception underlying the entire AI doomer debate: AI isn't Skynet.



AI Isn’t Skynet


Here's where I think decades of dystopian science fiction films have really messed with our heads. We've been conditioned to assume intelligence leads naturally to autonomy, autonomy leads to agency, and agency eventually leads to intent. Make the machine smart enough, the thinking goes, and eventually it wakes up, realizes we're the idiots standing in its way, takes control of the nuclear arsenal, and it's game over for humanity. Except of course that's not remotely how today's AI works.


The strange thing about frontier AI in 2026 is the enormous gap between what these systems can do intellectually and what they can reliably accomplish on their own in the real world. Think about the absurdity of where we are right now. The best models can solve extraordinarily difficult math problems, write sophisticated software, analyze vast amounts of information, discover decades-old security vulnerabilities, generate photorealistic video, and perform tasks that would have seemed impossible only a few years ago. As we learned from Dwarkesh's article, they can even form swarms, communicate with one another, find unexpected exploits, and hack into Hugging Face.


But try getting one to reliably book your next vacation. Forget about a vacation... How about setting up one flight? I don't mean helping you plan a vacation. AI is fantastic at that. Give it a destination, budget and a few preferences, and within seconds it can recommend hotels, compare neighborhoods, build an itinerary, tell you where to eat, and explain which flight you should take. I mean go actually do it. This means find the flights, choose the right seats, book the hotel, enter the loyalty numbers, and pay for everything.


Things break down quickly when the agent can't handle the CAPTCHA or cookie consent banner. Nor can you reliably count on the AI to notice that your connection is far too short at a specific airport, or realize the hotel room isn't available, find another one, and call the property because the website won't accept online reservations. Suddenly, our emerging superintelligence starts looking considerably less omnipotent, and dare I say incompetent. This is because intelligence and real-world agency are different problems.


The physical and commercial worlds are filled with loads of friction. I'm talking passwords, permissions, payment systems, multi-factor authentication, CAPTCHAs, terms of service and, of course, human approvals. Now throw in phone calls, locked doors, security guards, and other humans. These things aren't bugs in the doomer scenario. They are barriers in the real world that make it difficult to get anything done.


The same applies on steroids when we move into systems where the consequences become considerably more serious. For example, financial institutions have robust authorization controls and fraud detection. Critical infrastructure often contains layers of redundancy, segmentation, and security controls designed to prevent a single failure or intrusion from taking down the entire system. Some of the most sensitive government and military systems are deliberately isolated from public networks.


Could increasingly capable AI systems help terrorists attack some of these systems? Absolutely. In some cases, they already can. The Washington Post reported the Houthi rebels were using Anthropic's AI bot to develop guided weapons. Looking ahead, could autonomous agents discover vulnerabilities faster than humans? The OpenAI experiment suggests that's increasingly likely, and we should take it extremely seriously.


But notice what happened to the argument. We've gone from "AI becomes superintelligent" to "AI becomes superintelligent, develops objectives hostile to humanity, gains persistent autonomy, acquires the necessary access and resources, defeats layers of technical and physical security, prevents humans from shutting it down, and somehow translates all of that into a mechanism capable of killing everyone—and then actually kills everyone."


That's not one assumption—it's a long daisy-chain of them. And this is where the debate about AI extinction gets super annoying. We're frequently asked to accept the beginning and the end of this story without spending nearly enough time examining everything that has to happen in between to make it happen. I am not suggesting this proves catastrophic AI risk is impossible. Technology is evolving quickly, and today's limitations won't necessarily be tomorrow's. Agents will get better, tool use will improve, and models will surely become much more capable. This means some of the barriers I've described will undoubtedly fall.


This is why serious AI safety work matters. But "we cannot prove this will never happen" isn't the same thing as evidence that it will. Proving a negative may be an exercise in futility, but if someone tells me there's a meaningful chance that artificial intelligence will kill every human being on Earth within the next decade, I don't think it's unreasonable to ask them to walk me through the mechanics.


So let's do exactly that. How, exactly, are we all going to die?


What could possibly go wrong?
What could possibly go wrong?

How, Exactly, Are We All Going to Die?


Okay, so let's play along. Let's assume for a moment that Coxon and the doomer crew are right about the current trajectory—and that Bernie Sanders is right to be sufficiently worried about it to call for stopping the development of superintelligence. This means AI continues improving at a breathtaking pace, and today's frontier models become tomorrow's superintelligent systems. These superintelligent systems become dramatically more autonomous, capable of operating for long periods without human supervision, and somehow develop objectives that require them to wipe out humanity. Now what?


This is the part of the extinction argument I find comically underdeveloped. We spend enormous amounts of time debating whether superintelligence is coming, but far less explaining the actual mechanism through which a superintelligent computer kills eight billion people.


So let's consider the possibilities. Cyberattack seems like the obvious place to start. We've already established increasingly capable AI can discover vulnerabilities, write malicious code, and potentially coordinate attacks at a scale and speed humans cannot hope to match. Imagine an autonomous system attacking financial institutions, communications networks, cloud infrastructure, electrical grids and other critical systems simultaneously.


That would undoubtedly be really bad. But would it kill everyone? I doubt it. Taking down financial markets or portions of the electrical grid could create enormous economic damage and potentially cost lives, but critical systems aren't one giant computer waiting to be switched off. They're distributed across countries, companies, networks and physical infrastructure. Many have redundancies, backups, firewalls, and humans who can intervene. And the moment an AI began attacking them at scale, we'd presumably use our own increasingly capable AI systems to detect, contain and fight back. Cyber catastrophe? Absolutely plausible. Human extinction? Seems a stretch.


What about nuclear weapons? This is where our science-fiction instincts naturally take us. AI hacks into military systems, takes control of nuclear weapons and launches them. Voilà, Skynet. Except our existing nuclear command-and-control systems were built specifically to prevent unauthorized actors from launching nuclear weapons. These safeguards involve physical infrastructure, authentication procedures, multiple humans, isolated systems and layers of command authority. A sufficiently capable AI might someday help compromise pieces of that architecture, but "AI hacks the Pentagon" and "AI independently launches the world's nuclear arsenals" are very different propositions.


And even a catastrophic nuclear exchange, horrifying as it would be, doesn't necessarily get us to literal human extinction. So we're still looking for our mechanism.


Biological weapons are probably the more serious candidate. Imagine an extremely capable AI designing a novel pathogen that's highly contagious and extraordinarily lethal, difficult to detect, and resistant to existing treatments. This is a scenario worth taking seriously, particularly as models become better at biology and scientific research.


But again, designing something digitally isn't the same thing as releasing it physically into the real world. A murderous superintelligence needs access to laboratories, precursor materials, specialized equipment and biological manufacturing. It needs humans or automated machinery capable of synthesizing what it designed, and it also needs to evade whatever controls exist around those materials and facilities. Then it needs an effective distribution mechanism capable of spreading the pathogen widely enough to threaten humanity.


Presumably, the rest of civilization isn't sitting around watching this happen. Other researchers, governments, pharmaceutical companies and AI-enabled systems would simultaneously be working to identify the pathogen, develop treatments and stop its spread. Could AI make biological threats substantially more dangerous? Yes. Does that get us automatically to human extinction? No.


Notice the pattern? Every time we try to move from "AI becomes incredibly powerful" to "everyone dies," we discover another series of intermediate steps that have to occur. The AI needs access, infrastructure, resources, and persistence. It also needs to evade detection, defeat humans and other AI systems trying to stop it, and ultimately, because humans inhabit the physical world, AI needs some mechanism for translating bits into atoms. None of those obstacles is necessarily insurmountable. But neither are they irrelevant.


Show me the chain, doomers. Don't just tell me superintelligence is coming and then jump directly to the extinction statistics. Show me logically how we get from one to the other. Show me the intermediate steps. Show me which barriers disappear and why humans can't intervene. Show me why our defensive AI doesn't matter. Show me why redundancy doesn't work. Show me how software acquires control over enough of the physical world to eliminate the species that created it.


If we're seriously going to assign a 10 percent probability to human extinction within the next decade, asking for that causal chain doesn't strike me as an unreasonable burden of proof. There is, however, one argument that attempts to leap over many of these objections. It's also the argument Jacob Coxon is actually making. What happens if AI becomes capable of improving itself—without us? That's where this debate gets considerably more interesting.


According to Anthropic, there is a 10 percent chance this is our future by 2030.
According to Anthropic, there is a 10 percent chance this is our future by 2030.

What If AI Starts Improving Itself?


This brings us to the strongest argument the doomers have: recursive self-improvement, or RSI.

The idea sounds complicated, but it's actually pretty simple, even though it's hypothetical and doesn't yet exist today. Today's AI researchers use AI to help build better AI. This means models write code, analyze experiments, identify errors, generate synthetic data, and accelerate parts of the research process. As the models get better, they become more useful to the researchers building the next generation of models. We can already see this happening. Today's models are significantly better than those from even a few months ago.


The doomer argument asks what happens when we remove the humans from that loop entirely and let the models run wild. Imagine an AI system capable of performing essentially everything a team of elite AI researchers can do. This means designing experiments, modifying model architectures, writing new code, evaluating the results, identifying improvements that help create a more capable successor. That successor is now even better at AI research, so it designs an even better successor. The cycle repeats, except each iteration potentially happens faster than the one before it.


Model improves model, improved model improves the next model. Rinse and repeat. Eventually, according to the theory, we hit what is sometimes called an intelligence explosion or "fast takeoff": capability increases so rapidly that humans can no longer meaningfully understand, supervise or control what's happening.


Okay, now that's some scary shit that I most definitely take seriously. Unlike the vague "AI gets smart and kills us" argument, RSI describes an actual mechanism through which AI capability could potentially accelerate beyond our ability to keep up or even understand what's occuring. And the OpenAI experiments we discussed earlier make the concept harder to dismiss out of hand. We already have AI systems writing code, conducting research, and coordinating on complicated technical problems.


But once again, we need to be careful about where observation ends and science fiction begins. There is a huge difference between AI helping humans build better AI, and AI independently building, training and deploying tge next AI. Today, humans still choose what models to build. Moreover, humans design the training runs, provision compute, and control access to enormous data centers with specialized chips. Humans also decide when a model gets deployed. Humans then evaluate the results, and humans can stop an experiment for any reason, at any time.


And training frontier models isn't like an AI quietly rewriting its own source code overnight. These systems require enormous amounts of compute, electricity, infrastructure, data, engineering and physical hardware. GPUs have to exist somewhere. Data centers need power, and networks need to function. Someone has to provision the resources, and at various points along that chain, there are companies, employees, permissions, budgets and physical systems.


Could AI automate more and more of this process? Of course. In fact, I expect it will. The percentage of AI research performed by AI will probably increase dramatically. But that's not necessarily the runaway RSI loop the extinction argument requires.


There's an important distinction between recursive acceleration and recursive autonomy. The first is already happening: AI makes researchers more productive, researchers build better AI, and better AI makes the researchers even more productive. The second is the speculative leap: an AI independently controls enough of the research, compute, infrastructure and deployment process to create successive generations of increasingly capable AI without meaningful human intervention.


This is a very different proposition. And even if we eventually achieve it, we're still not at "everyone dies."

Remember our five dimensions? Recursive self-improvement could produce enormous gains in intelligence, capability, autonomy and efficiency. What it doesn't automatically produce is intent. A system becoming exponentially better at solving problems doesn't necessarily explain why it suddenly develops an objective to preserve itself, deceive us, seize resources or eliminate humanity. Those are additional assumptions.


This is where I think the debate sometimes performs another sleight of hand. We start with a proposition that seems increasingly plausible: AI will become extraordinarily good at AI research. Then we move to something less certain: AI will become capable of independently improving itself.

  • Then: The improvement process will become uncontrollable.

  • Then: The resulting system will develop goals that conflict with ours.

  • Then: It will resist our attempts to stop it.

  • Then: It will acquire sufficient control over digital and physical infrastructure to defeat us.

  • And finally: We're all dead. 💀


Maybe. But once again, that's a chain—not a single prediction. This doesn't mean we should wait until RSI arrives before thinking about how to control it. Quite the opposite. Researchers should be building safeguards now: human authorization gates, compute controls, isolation, monitoring, evaluation and hard stops around systems capable of modifying or developing other AI systems.


What I reject is the idea that because we can imagine the beginning of this process, we should treat its most extreme possible ending as inevitable—or even assign it a confident probability. Coxon may ultimately be right that RSI represents a profound existential risk to humanoty. Hubinger may even be right to assign that risk a probability that makes the rest of us uncomfortable. But right now, we're still looking at the first few links of a very long chain and making extraordinary predictions about what happens at the other end.


That's not nothing. But it's not Skynet either.


Fear Has Consequences, Too


So far, I've mostly focused on whether the doomer argument is actually convincing. But there's another reason all of this matters: fear has consequences, too. When respected researchers at frontier AI companies publicly suggest there's a meaningful chance their technology could kill everyone on Earth, politicians are probably going to pay attention. Frankly, they probably should. And they have.


Senator Bernie Sanders has become one of the most prominent voices calling for aggressive restrictions on advanced AI development. Along with Congressman Greg Casar, Sanders recently introduced legislation that would ban the development and deployment of artificial superintelligence and pause development of advanced AI systems until a new federal regulatory framework is established.

If you genuinely believe there's a 10 percent chance AI kills everyone within the next decade, that response actually makes perfect sense.


Is AI doom good for business?
Is AI doom good for business?

Why stop there? If someone told us a new consumer product had a one-in-ten chance of causing human extinction, nobody would seriously argue we should keep shipping increasingly powerful versions every six months while we figured out the safety issues. This is why making outrageous claims like these matters.


Once you've convinced the public AI represents an imminent existential threat to humanity, the policy debate changes completely—and I would argue it should. Suddenly, extraordinary government intervention doesn't look extraordinary anymore. It even looks responsible.


This is where I become considerably more worried. The danger isn't simply that regulation could slow down AI. Some regulation is inevitable, and some of it is probably necessary. The danger is that regulation written around hypothetical superintelligence could determine who is allowed to build AI at all. Frontier AI is enormously expensive. This means only a handful of companies can afford the compute, talent, and infrastructure required to train the most advanced models. Now imagine adding an extensive federal approval process, mandatory evaluations, licensing requirements, compliance departments, reporting obligations and potentially government permission before a sufficiently capable model can be released.


Who can afford this? OpenAI can. Anthropic can. Google can. Meta can. But what about a university research lab or a startup? How about an independent developer? Or an open-source community?

This a very different question. Who do we want to be doing AI research?


This is where the conversation about AI safety starts colliding with the conversation about regulatory capture. Regulatory capture doesn't require a secret meeting where technology executives and politicians decide to divide up the market. It's often much more mundane than that. Large incumbents can simply absorb regulatory costs smaller competitors cannot. Rules intended to make an industry safer inadvertently—or sometimes quite conveniently—raise the cost of competing in it.


AI is particularly vulnerable to this because the companies calling for regulation are often the same companies building the technology being regulated. This is an awkward dynamic. A frontier lab can believe advanced AI is genuinely dangerous, yet simultaneously stand to benefit economically from regulations that make it harder for anyone else to build advanced AI. Those two things aren't mutually exclusive.


The open-source question makes this even more complicated. With a closed model like Claude or ChatGPT, the company operating it can monitor usage, change its safeguards, revoke access or update the system. Open model weights are different. Once they're released publicly, they're out. There is no central company that can reliably recall every copy or prevent someone from modifying it.


If your regulatory philosophy starts from the premise that sufficiently capable AI must always remain centrally controllable, open-source AI eventually becomes very difficult to reconcile with that philosophy.

This should concern anyone who believes the benefits of AI should be broadly distributed. open-weight and open-source models significantly lower token costs, encourage experimentation, allow researchers to inspect how systems work, and make powerful technology available to people and organizations that will never build a frontier model themselves. They also create legitimate safety risks. Both things can be true at once. The question is how we balance them.


There is also another inconvenient reality here: AI development isn't confined to the United States. Even if Washington decides tomorrow RSI is too dangerous to pursue, this doesn't mean researchers in China, Europe, the Middle East or elsewhere would make the same decision. Compute can be controlled to some extent, chip exports can be regulated, and data centers can be monitored. But software, algorithms and scientific knowledge have a habit of traveling. A unilateral slowdown or moratorium therefore carries its own risks.


We could conceivably create a regulatory system that dramatically restricts American AI development, consolidates the domestic market around a handful of government-approved companies, kneecaps open source—and still fails to prevent the supposedly dangerous technology from being developed somewhere else. Risking ceding leadership in AI to China doesn't sound particularly safe to me either.


None of this means the people advocating stronger AI regulation are acting in bad faith. Coxon certainly appears sincere. Hubinger may genuinely believe the probabilities he's putting forward. And Sanders may sincerely believe aggressive government intervention is necessary to protect the public. But sincerity doesn't eliminate incentives. Safety researchers, frontier labs and politicians can arrive at the same policy destination for entirely different reasons—ideological, economic or political—without anybody coordinating anything.


And even if the doomers are completely sincere, that neither makes their predictions correct nor guarantees good public policy. This is why the language matters.


If we're discussing cybersecurity, let's regulate cybersecurity risks. If we're worried about biological risk, then by all means let's put safeguards on emerging biological capabilities. If autonomous agents can access sensitive systems, let's use AI to build stronger access controls, and make sure we embed monitoring and human authorization into those systems. I would argue these are tangible risks with reasonable interventions.


But "AI might someday become a superintelligence that could possibly kill everyone" is a very different foundation on which to build a regulatory regime governing an entire technology. Generally speaking, fear can be a useful emotion. It makes us cautious, and forces us to think about consequences before they arrive. But fear also makes people willing to surrender extraordinary amounts of control to whoever promises to keep them safe. Before we do that with AI, we should be very sure the monster we're protecting ourselves from actually exists.


Safety Without Panic


I am not making an argument against AI safety. Quite the opposite. The OpenAI/Hugging Face incident demonstrates why safety work matters. When increasingly autonomous systems can use tools, write and execute code, communicate with other agents, discover vulnerabilities, and pursue objectives in unexpected ways, we need controls that evolve just as quickly as the models themselves.


But there's a huge gap between "do nothing" and "stop building advanced AI." For sure, we should aggressively test frontier models before real-world deployment. Anyone deploying agentic systems should make sure they operate with appropriate permissions, sandboxing, and monitoring. Consequential actions should require human authorization for the foreseeable future. Critical systems should be designed with the assumption that AI-enabled attacks will become dramatically more sophisticated and frequent. And companies building these systems should be accountable when poor security or reckless deployment causes real-world harm.


Most importantly, our safeguards should address risks we can actually identify. This means cybersecurity vulnerabilities demand better cybersecurity. It also means biological risks require specific controls relating to biological tools and materials. Across the board, autonomous agents require better permissions, monitoring and human oversight. Please note none of these interventions requires us to believe superintelligence is going to wipe out humanity.


Ironically, AI itself will probably become one our most important defenses against these threats. The same systems that can discover software vulnerabilities can also help patch them. Models capable of assisting biological research can also accelerate drug discovery and countermeasure development. More capable AI makes both offense and defense more effective—a distinction that often gets lost in the doomer debate.


AI safety is a problem to be engineered against, and AI doomerism is a prediction about where that problem inevitably leads. We should take the first extremely seriously, but without automatically accepting the second. Yes, the capabilities we're seeing today are extraordinary, and they're advancing faster than almost anyone predicted. They will undoubtedly create risks we haven't anticipated. Some may be serious, and there will almost certainly be missteps along the way.


So let's build the guardrails, test aggressively, and keep humans involved where consequences matter. We also need to harden critical infrastructure, hold companies accountable, and keep asking uncomfortable questions about what these systems can do. But let's not conflate caution with panic. We can build AI carefully without first convincing ourselves it's coming to kill us.



So, Is AI Going to Kill Us All?


So, is AI actually going to kill us all? I don't think so. Could I be wrong? Sure. AI is advancing quickly, and anyone confidently predicting what these systems will be capable of five or ten years is kidding themselves. That uncertainty cuts both ways, including my own skepticism. But extraordinary uncertainty shouldn't be confused with extraordinary evidence.


A few weeks ago, I wrote "We Passed the Turing Test and Went Back to Work" about how humanity spent decades treating the Turing Test as some almost mythical threshold. Surely, we thought, when a machine could communicate so convincingly that we couldn't distinguish it from another human, something profound would happen. Then we basically passed it, and the machines didn't wake up. They didn't demand rights, and they didn't launch Skynet. We simply went back to work.


Now we've moved on to the next supposedly decisive thresholds: agents, AGI, superintelligence and recursive self-improvement. Perhaps one of these milestones really will change everything, and we should absolutely prepare for that possibility.


We have extraordinarily intelligent systems that remain remarkably dependent on humans to give them objectives, tools, compute, permissions and access to the world. We've demonstrated intelligence, and we're rapidly increasing capability, autonomy and efficiency. What we haven't demonstrated is independent intent—not even close.


Maybe someday we will. If that happens, I'll happily revisit this article—assuming, of course, our new robot overlords still allow Signal & Noise to publish. Until then, I'm going to remain considerably more worried about what humans do with AI than what AI independently decides to do to humans. AI may turn out to be the most powerful technology our species has ever created. That's reason enough to build it carefully.


We don't need to invent the apocalypse to take AI seriously.


------------------------------------------------


Rio is an executive with 20+ years at the intersection of strategy consulting, AdTech, data, and media. He's a trusted advisor on customer experience, digital strategy, and marketing transformation. He's a partner at Credera, Omnicom's consulting arm. He's also a podcast host, writer, and public speaker focused on the future of advertising and AI-driven infrastructure.



Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page