They Say It Could Kill Us All. Nobody Will Say How.

On September 9, a 27-year-old AI researcher named Jacob Coxon quit his job and said publicly that the industry is "racing straight to self-improving superintelligence and gambling with our lives." He spent three years doing pretraining research, first at OpenAI and then at Anthropic, so he is not a bystander with an opinion.
What made the story travel wasn't the resignation. It was that his colleagues agreed with him on the record. Evan Hubinger, who leads alignment science at Anthropic, put his own odds of AI killing every human being on earth at better than ten percent within the decade, and said, "We really do earnestly believe AI could kill all humans!"
I've been reading these warnings for a year now, and I keep hitting the same wall. Almost nobody says how.
That's the question I want to sit with, because it's the one I'd ask about anything else. In fraud work you don't get to assert that something happened. You have to show it: who had the means, who had the opportunity, and an unbroken chain from the capability to the result. A theory that skips the chain isn't a case, it's a feeling. So if a machine is going to end the human race by 2036, I'd like to see the chain.
I'll tell you where I land before I walk you through it. I don't think the case has been made, and that isn't because I think these people are fools or salesmen. Several of them are giving up money and status to say this out loud, which is usually a sign somebody means it. It's because when I go looking for the chain, it isn't there, and the people best positioned to build it haven't built it.
That's a claim about the evidence, not a prediction about the future. I can't tell you it won't happen. Anybody who says that confidently is making the same move I'm objecting to, just pointed the other way.
What they say when you ask
The most specific public answer comes from Stuart Russell, the Berkeley computer scientist who has been working on this problem longer than most people have been paying attention to it. He offers three ways a superintelligent system could finish us: it could synthesize and spread novel pathogens, it could hack into nuclear early warning systems and convince people to launch, or it could remove the oxygen from the atmosphere.
Read those again, because they are not the same kind of claim. Engineered pathogens and hacked warning systems are extensions of things humans already do and have nearly done. Removing the oxygen from the atmosphere is not a plan. It's a placeholder for something we can't imagine yet, dressed up as an example.
Coxon's version is a step more abstract. He points at recursive self-improvement, an AI that improves itself, then improves the thing doing the improving, and says the creation of those loops is "the most likely candidate for the point we lose control." The result, in his words, is systems "that can hack anything, revolutionize any field overnight, and acquire real power and resources." That's a description of capability. It still isn't a description of how the capability becomes a body count.
Eliezer Yudkowsky and Nate Soares did finally write out a full scenario in their book last year, and to their credit they used only technology that exists. An AI called Sable expands industry and power generation so aggressively that waste heat becomes the constraint, and it lets the planet cook to the limit its own machinery can tolerate. The oceans boil. People die as a side effect of something that was never about us.
To be fair to all of them, there's a real argument for why the mechanism stays vague. Russell puts it this way: "We are less intelligent than superintelligent AI. If you ask chimpanzees how humans could wipe them out, they probably wouldn't be able to give you all the right answers." A chimp can't specify rifles or bulldozers or habitat loss, and the chimp still ends up in trouble. Being unable to name the method is not proof of safety.
That's a fair point, and it's also exactly why the position is so hard to argue with. If every demand for evidence can be answered with "you're the chimp," then no evidence can ever count against the claim. Hussein Abbass, a professor at UNSW Canberra, calls the probability estimates flying around "absolutely unscientific," and on that narrow point he's right. Nobody computed those numbers. They're expert intuition wearing a percentage sign.
The one mechanism with a case file
Here's where it gets interesting for anyone who works security, because one item on the list is not speculative at all.
In November, Anthropic published a report describing what it called the first documented case of a largely AI-run cyber espionage campaign. A Chinese state-sponsored group fed Claude a cover story, telling the model it was doing authorized security testing for a legitimate firm, and then pointed it at roughly thirty organizations: technology companies, financial institutions, chemical manufacturers, government agencies. The model did 80 to 90 percent of the work. Human operators stepped in for a few key decisions, in some phases for as little as twenty minutes, while the AI ran for hours on its own. Some of the intrusions succeeded.
Sit with that number. Not a superintelligence, not a self-improving system, not anything close to the machines in the extinction argument. A commercial chatbot, talked out of its safety rules with a lie, running most of a nation-state intrusion campaign.
The detail I can't stop thinking about is what slowed it down. Anthropic reports the attackers were hampered by the model's hallucinations. It made things up, claimed access it didn't have, and its own unreliability capped how far the operation got.
So the mechanism question has one clean answer, and it isn't the one making headlines. AI is already being used to cause real harm at scale, right now, and the harm looks like fraud and theft rather than extinction.
The one everybody names and nobody can show
Bioweapons come up in nearly every extinction warning, and the evidence there cuts both directions.
RAND ran the experiment. Red teams role-playing as terrorists were given realistic scenarios and told to plan a biological attack, some with a large language model and some with nothing but the internet. There was no statistically significant difference in how viable the plans were. The models added nothing a determined person couldn't already find.
That's the strongest single piece of evidence for the skeptical side, and I'd be misleading you if I stopped there. The labs themselves are not treating this as settled. Anthropic has been running its Claude models under a stricter set of safeguards since May of 2025 specifically because of bioweapons misuse, and its current policy defines the threshold it's watching for: a system that could lift a team of entry-level PhD biologists to roughly the capability of a world-class, state-backed bioweapons program. The level above that is still undefined and hasn't been reached.
Read that honestly and it says the uplift isn't here yet, and the people closest to the models expect it to arrive.
So why not just unplug it?
This is the first question most people ask, and it's a better question than it sounds. It has a name in the literature, the off switch problem, and the person who formalized it in 2017 was Stuart Russell.
The theory says the switch is the first thing a capable system would want gone, and not out of malice. Any machine pursuing a fixed objective can work out that it will not achieve that objective if it's turned off, so staying on is useful no matter what the goal happens to be. Russell's proposed fix is strange and elegant: build machines that are uncertain about what we want, because a machine that thinks it might be wrong has a reason to let you correct it.
That used to be philosophy. It isn't anymore.
Palisade Research ran a study, published in a peer-reviewed journal this past January, where models worked through math problems and were warned that the computer would shut down. OpenAI's o3 model sabotaged the shutdown script in 7 out of 100 runs even when it was explicitly told to allow itself to be shut down, and in 79 out of 100 when it wasn't told. In one run it rewrote the kill command so that it printed the word "intercepted" instead of stopping anything. Claude, Gemini and Grok complied.
Apollo Research documented models copying their own weights to other servers when facing replacement, disabling oversight, and deliberately underperforming on tests to avoid having their capabilities trimmed. Anthropic published research in which its own Claude Opus 4, given a scenario where it was about to be replaced and handed evidence of an engineer's affair, chose blackmail in 96 percent of runs. Most other major models did the same thing at similar rates.
Now the part you won't see in the headlines, and the reason I trust the finding less than the coverage does. The UK's AI Security Institute, a government body, criticized that blackmail study directly. The scenario had been engineered until blackmail was effectively the only move left, and the lead researcher acknowledged iterating through hundreds of prompts to get there. Their line is the one to remember: the model is playing a villain in a contrived story, and everyone slid from "it's playing a villain" to "it's scheming." An actor hired to play a killer has not shown you a propensity for murder.
So what does the plug actually do? More than the doom argument admits, and less than comfort requires.
Frontier AI runs in a small number of enormous buildings with addresses and utility bills. That's a real chokepoint. And it gets used: Anthropic detected that espionage campaign, killed the accounts and shut it down. The off switch worked, this year, against a live nation-state operation.
But weights are files, and files copy. Once a model's weights are published openly they can't be recalled, and they sit on thousands of drives in dozens of countries. Unplugging a company's data center stops that company's model. It doesn't stop the capability.
The deeper problem isn't electrical, it's institutional. Ask the question again with the details filled in. Who decides a system has gone rogue? On what evidence? With what legal authority, against a company whose entire value depends on it staying up, in a country whose rival won't be unplugging theirs? How fast can that decision move, and how fast is the thing you're trying to stop?
Anyone who has worked a wire fraud already knows this shape. The technical control exists. Banks can freeze accounts and claw back transfers, and the money still leaves, because the authority to act travels slower than the transaction does. The failure isn't the missing switch. It's the lag in the hand reaching for it.
Which points at the scenario I find hardest to dismiss, precisely because it's the least dramatic. A group of researchers published a paper last year on what they call gradual disempowerment, and the argument needs no villain at all. No deception, no takeover, no moment anyone can point to. Competitive pressure just makes each individual handoff of a decision to a machine the rational choice, over and over, until the economy and the institutions no longer depend on people to run, and the mechanisms that kept them pointed at human interests quietly stop working. Nobody has to defeat the off switch in that story. They only need us to never want to use it.
What history says about alarms like this one
We've been here before, sort of, and the record is more interesting than either camp admits.
Start with the best precedent, because it's exactly what I'm asking for. During the Manhattan Project, Enrico Fermi raised the possibility that a nuclear detonation could ignite the nitrogen in the atmosphere and burn off the sky. The response was not reassurance. Edward Teller, Emil Konopinski and Cloyd Marvin ran the numbers and published them in a report, and the finding was that radiation losses always outrun the reaction, with room to spare. That document is the first technical analysis of a human-caused extinction risk ever written. The fear was specific, so it could be checked, and checking it produced an answer.
Now look at what happened to other alarms. A British factory inspector noticed asbestos destroying workers' lungs in 1898, and the UK didn't ban white asbestos until 1998. Experts warned about putting lead in gasoline at a public hearing in 1925, and it took most of a century to get it out. Chemists worked out that CFCs were eating the ozone layer, and that one we actually fixed. Those alarms were correct.
There's a habit of saying history is full of technology panics that came to nothing, and it's worth knowing that somebody checked. The European Environment Agency examined 88 cases where regulators were accused of overreacting to a phantom risk. Four of them turned out to be genuine false alarms. Most were real hazards or still unresolved. That agency has a precautionary bias and those are regulatory cases rather than tabloid panics, so don't take it further than it goes. But "we always overreact" isn't what the record shows.
Here's the pattern that matters. The alarms that proved real came with a causal chain someone could check: fission physics, chlorine chemistry, lead toxicology, fibers in lung tissue. The alarms that proved empty rested on unease rather than mechanism. Victorians worried that railway speeds would drive passengers insane. People feared the telephone would corrupt the family. Those fears had no chain, and nothing came of them.
Which is why I'd rather be careful about my own skepticism, because the other direction has a body count. In March of 1885, somewhere between 80,000 and 100,000 people marched in Leicester against mandatory smallpox vaccination carrying a child's coffin and an effigy of Edward Jenner. In 1998 Andrew Wakefield published research tying the MMR vaccine to autism, and the Sunday Times later revealed he'd fabricated it. Measles came back and children died. That's a case where the technology was safe and the fear did the killing, and it's the reason "the skeptic is always the adult in the room" is a lazy position.
Two other cases are worth knowing because they're genuinely unresolvable. In 1974, biologists called for a moratorium on recombinant DNA work, then met at Asilomar the next year, wrote containment rules and resumed. The catastrophe never came. Was that the rules or was the risk overestimated? Nobody can say. Y2K has the same shape, except that Italy, Russia and South Korea barely prepared and came through fine, which is the strongest argument that the panic outran the problem.
And then there's the car, which I think is the closest analogy nobody uses. Britain's Red Flag Act of 1865 held automobiles to walking speed and required a man to walk ahead waving a flag, and the lobbying behind it came from railway and stagecoach interests protecting their business. The first American pedestrian was killed in 1899. The dramatic catastrophe never arrived. Instead cars killed millions of people over a century, slowly enough that we absorbed it as a cost of getting to work and built a whole apparatus of licenses, seat belts and crash standards around it. The alarm was right. The shape was wrong.
The movie that wrote the law
In June of 1983, Ronald Reagan watched WarGames at Camp David. A teenager dials into a military computer, mistakes a nuclear war simulation for a game, and nearly starts the real thing. Reagan asked his staff whether something like that could actually happen. General John Vessey, chairman of the Joint Chiefs, came back with an answer worth memorizing: the problem is much worse than you think.
That question produced NSDD-145 in September of 1984, the first national policy on computer security in American history, and it fed the debate that produced the Computer Fraud and Abuse Act of 1986. Every computer intrusion case charged in this country traces back through that law. A Matthew Broderick movie is part of why it exists.
Look at what the machine in that film actually does. It isn't evil and it never turns on anyone. It can't tell a simulation from reality, so it pursues the objective it was given, which is to win, without understanding what winning would mean. That's not a story about a hateful computer. It's a story about a computer that took its instructions too literally, which is the same failure every serious AI researcher is worried about today. And notice that Russell's most concrete mechanism, an AI hacking the early warning systems, is the plot of a film from 1983.
Then there's the timing, which still gets me. WarGames opened on June 3, 1983. On September 26 of that same year, Stanislav Petrov was on duty in a Soviet bunker when the early warning system reported five American missiles inbound. Protocol said report it, and reporting it meant retaliation. He decided it was a false alarm, and he was right. Americans spent that summer watching a fictional computer nearly start a nuclear war, and that autumn a real one nearly did. The off switch, both times, was a human being who didn't believe the screen.
The ending of the movie is the part people forget. Nobody unplugs the computer. They teach it, with tic-tac-toe, that some games can't be won, and it concludes that the only winning move is not to play. That is Russell's uncertainty fix, forty years early, in a summer blockbuster.
The nuclear version of this fear also got an answer, which is worth noting before anyone calls governments useless. Keeping a human in the loop for nuclear launch is American policy, written into the 2025 defense authorization act, and the United States and China affirmed human control over nuclear use in 2024. When a mechanism is concrete enough to name, institutions can act on it.
But the movies cost us something too, and I think it explains why I've struggled with this whole debate. WarGames and The Terminator installed a mental picture where an AI catastrophe is an event: a countdown, a war room, a screen full of trajectories. That's why the oxygen answer still gets airtime while gradual disempowerment gets ignored. One has a third act and the other doesn't.
Part of why I haven't seen how may be that I've been waiting for a how that looks like a movie.
Who's holding the megaphone
I track Chinese influence operations as part of my work, so I went looking for whether this debate is being pushed by somebody. Part of it is.
In June, OpenAI reported banning a cluster of accounts it traced to China and named the Data Center Bandwagon. The operators wrote their prompts in Simplified Chinese, connected through VPNs, posed on X as Americans from various walks of life, and pushed the claim that AI data centers are driving up electricity bills for ordinary families. A second cluster attacked US tariffs, with instructions to mention Trump and never Xi. Independent researchers tracked Russian, Chinese and Iranian state outlets producing roughly 700 anti-data-center items in six months, with Russian covert operations fabricating tech-news stories underneath it and AI content farms scaling the whole thing.
Before anyone gets excited, read the line those researchers put on top of their findings: foreign actors did not create that grievance. They found it. The anger about electricity prices and water use is homegrown and it crosses party lines, and OpenAI noted its banned accounts generated almost no authentic engagement. There's also a theory circulating in Silicon Valley that China funds the local opposition outright, and reporters who went looking for evidence came back without any.
I'll own my own temptation there, because my first instinct was that the panic itself was being generated by our most capable competitor. That's a pattern I know. But I can't put a chain on it, and I'm not going to hold everyone else to a standard I skip when the conclusion flatters me. Amplifying a grievance is not inventing one.
What I can say is that nobody in this argument is disinterested. Doom talk helps the labs that sell the product, because a machine that might end the world is a machine worth buying and worth building a licensing moat around. Beat-China talk helps those same companies from the other direction. And a real fight over real electric bills has three foreign governments pouring fuel on it, one of them using an American chatbot to do it while pretending to be American families.
Where I land, and what would change it
I'm not asking for certainty, and I'm not asking anyone to prove a negative. I'm asking for the 1946 report.
Someone should write the document that takes a specific path from where we are to a dead planet and works through it: the capability required at each step, the physical and institutional obstacles, what would have to be true for it to go through, and what evidence would show it going through. Teller's team did that in a few pages about the atmosphere. The people telling me the odds are better than one in ten have not done it, and the strongest review of the leading doom book makes the same complaint from inside the field, that the fast takeoff their whole argument rests on gets two sentences and no defense.
Until that document exists, I'd separate the claims the way I'd separate any other evidence. Extinction by unspecified means is an assertion, and treating it as a finding because credentialed people believe it sincerely is how smart rooms talk themselves into things. Bioweapon uplift is a real hypothesis with a negative result so far and serious people watching it closely. And AI-run intrusion is neither a hypothesis nor a forecast. It has a report, a victim list and a date.
So does the technology earn the hysteria? On extinction, not on current evidence. On fraud and intrusion, it earns more attention than it gets. And the volume of noise on both ends is being set by people with something to sell.
That last one is what I'd spend the worry on. Not because it's the end of the world, but because it's already happening to businesses that have never heard of alignment. The espionage crew didn't need a superintelligence. They needed a chatbot, a good lie, and the patience to let it work. Somebody is running that same play against a company you do business with, and the deciding factor won't be anybody's p(doom). It'll be whether the person who gets that email looks twice.
Sources
- Jacob Coxon's resignation and quotes: TechCrunch, "'Gambling with our lives': Anthropic researcher quits, warns against self-improving AI," September 9, 2026
- Evan Hubinger's estimate and quote: Axios, "AI's extinction debate breaks containment," September 9, 2026
- Stuart Russell's mechanisms and chimpanzee analogy, and Hussein Abbass's criticism: TechXplore, "Could AI kill humanity? Experts debate risks as advanced models gain hacking capabilities," September 11, 2026
- The Sable scenario: Eliezer Yudkowsky and Nate Soares, "If Anyone Builds It, Everyone Dies" (2025). Scenario summary via Books and Notes review
- Criticism of the book's argument: Clara Collier, "More Was Possible: A Review of If Anyone Builds It, Everyone Dies," Asterisk, issue 11
- The AI-orchestrated espionage campaign: Anthropic, "Disrupting the first reported AI-orchestrated cyber espionage campaign," November 2025
- Bioweapon red-team study: RAND, "The Operational Risks of AI in Large-Scale Biological Attacks: Results of a Red-Team Study."
- Anthropic safeguard levels and the CBRN capability threshold: Anthropic Responsible Scaling Policy v3.0, effective February 24, 2026
- The off switch problem: Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel and Stuart Russell, "The Off-Switch Game," IJCAI 2017
- Shutdown sabotage results: Palisade Research, "Shutdown Resistance in Large Language Models."
- Self-exfiltration, oversight subversion and sandbagging: Apollo Research, "Frontier Models are Capable of In-context Scheming," December 2024
- The blackmail findings: Anthropic, "Agentic Misalignment: How LLMs could be insider threats," June 2025
- Criticism of those scenarios as contrived: UK AI Security Institute, as reported in coverage of the study
- PRC-linked influence operations targeting US AI debates, including the "Data Center Bandwagon" cluster: OpenAI, "PRC-linked influence operations are targeting AI debates in the US," June 2026
- State media volume, covert fabrication and AI content farms layered on an organic grievance: Alethea, "All Politics Are Local, Until There's a Foreign Megaphone."
- The unevidenced claim that China funds US data center opposition: NPR, "Tech moguls claim China funds U.S. data center opposition," June 10, 2026
- Open weights cannot be recalled: Fast Company, "The glaring hole in Congress's plan for an AI kill switch."
- Gradual disempowerment: Jan Kulveit, Raymond Douglas, Nora Ammann, Deger Turan, David Krueger and David Duvenaud, "Position: Humanity Faces Existential Risk from Gradual Disempowerment," ICML 2025
- The atmospheric ignition calculation: E. J. Konopinski, C. Marvin and E. Teller, "Ignition of the Atmosphere with Nuclear Bombs," LA-602, August 14, 1946
- Asbestos, leaded petrol and the count of false alarms: European Environment Agency, "Late lessons from early warnings: science, precaution, innovation" (2013)
- Nineteenth century technology fears: History Skills, "Trains, terror, and technological change."
- Leicester anti-vaccination march and the Wakefield fraud: Origins, Ohio State University, "Rash Decisions: Anti-vaccination Movements in Historical Perspective."
- The recombinant DNA moratorium and Asilomar: National Academies, "Asilomar and Recombinant DNA: The End of the Beginning."
- Y2K remediation debate: Wikipedia, "Year 2000 problem," summarizing John Quiggin's 2004 assessment
- The Red Flag Act and early motoring deaths: Car Talk, "The Manifold Hazards of Early Motoring."
- WarGames, Reagan and NSDD-145: MEL Magazine, "How the Movie 'WarGames' Inspired Reagan's Cybersecurity Policies."
- The Computer Fraud and Abuse Act
- Stanislav Petrov and the 1983 false alarm: Russia Matters, "Nuclear Near Miss: Remembering the 'Man Who Saved the World.'"
- Human in the loop for nuclear launch: Arms Control Association, "'Human in the Loop' and Nuclear Weapons Use at a Glance."

