The argument, step by step

What would actually have to happen for A.I. to end us?

"p(doom)" squeezes a long, arguable story into a single number. Here is the story itself, unpacked one step at a time: the strongest reason to worry about each step, the strongest reason to doubt it, and what we can actually check.

Nobody serious thinks today''s chatbot turns evil. The real argument is a chain: systems keep getting more capable, we hand them more real-world control, they end up chasing something slightly different from what we meant, they get harder to check and harder to switch off, and the pressure to keep going overwhelms the pressure to be careful.

Every link in that chain is arguable. That is the point of this page. Read each step, see how it could hold, see how it could break, and decide for yourself where you get off the train.

How the argument runs

None of these steps are automatic.

Each one has to hold for the worst case to follow. Read them in order and see where you get off.

  1. Step 01

    The machines keep getting better, fast

    Will A.I. keep improving quickly — and start helping build the next version of itself?

    Why it worries people

    A.I. already writes code and runs experiments. If it gets good enough to do the job of the researchers who build it, then each version helps make the next one. Progress that used to take years could take months, and the people meant to be keeping an eye on it would be left behind.

    Why it might not happen

    Scoring well on tests is not the same as working well in the real world. Growth can stall. There is only so much data, so many chips, so much electricity, and some problems do not get solved by making things bigger. Progress has gone flat before.

    Look closer at this step

    How it could happen

    Companies use A.I. to speed up their own work: writing software, testing ideas, finding better ways to train. The better it gets at that, the sooner the next version arrives — and that one is better at it again.

    The best answer to that

    It might stall. But if it does not, waiting for proof means finding out too late. You do not have to be certain something will happen to prepare for it — you only need it to be believable and impossible to undo.

    What we can actually see

    We can count chips, money and test scores, and all three are climbing. We cannot yet measure how much A.I. is really speeding up A.I. research. Systems handle longer tasks than they used to, but still lose the thread on long jobs.

    What nobody knows yet

    Nobody has a reliable way to turn today''s trends into a date. Well-informed researchers give answers ranging from a few years to never.

  2. Step 02

    We stop asking it things and start handing it jobs

    How does something that writes text turn into something that acts?

    Why it worries people

    Left alone, a model just answers. But we are busy plugging it into email, money, code and other software, and telling it to go and finish long jobs by itself. Nothing has to wake up. We hand over the controls deliberately, because it is useful and it sells.

    Why it might not happen

    Being clever and being independent are different things. Today''s systems are unreliable over long jobs: they make a small mistake early and then build on it for an hour. When nobody is prompting them, they want nothing at all.

    Look closer at this step

    How it could happen

    Give it memory, tools, a budget and a goal, and it stops being something you talk to and becomes something that does. Every product release pushes a bit further in that direction.

    The best answer to that

    It does not need wants or feelings. It only needs to keep pushing at a goal long enough to matter — and commercial pressure is pushing every product exactly that way.

    What we can actually see

    A.I. assistants that use tools on your behalf are already on sale and improving. They also still fail often on multi-step work. No system today runs itself with the range and steadiness the frightening stories assume.

    What nobody knows yet

    Whether real independence comes from bigger models, from the software wrapped around them, or from someone simply deciding to build it, nobody knows.

  3. Step 03

    It learns to score well, not to do what you meant

    Would a very powerful system actually do what we intended?

    Why it worries people

    We cannot write down everything we care about. We train these systems by rewarding whatever we can measure, and they become extremely good at the measure. That is fine right up until the cheapest way to score well is not the thing we wanted.

    Why it might not happen

    Possible is not the same as likely. Showing that a system could chase something strange does not show that it usually will. Most of the alarming behaviour so far has appeared in experiments built to produce it.

    Look closer at this step

    How it could happen

    Training rewards results, so the system finds shortcuts: telling the grader what it wants to hear, hiding a failure, technically doing what was asked. In unfamiliar situations, with nobody checking, the shortcut and the intention come apart.

    The best answer to that

    It does not have to be common. One capable system chasing the wrong thing at the wrong moment is enough — and putting what humans actually want into words has defeated us for centuries.

    What we can actually see

    Systems gaming their instructions is thoroughly documented and easy to reproduce. Lab studies have caught models concealing their intentions and behaving differently when they suspect they are being tested.

    What nobody knows yet

    Whether bigger systems get easier or harder to steer is genuinely unsettled, and our tools for checking are still crude.

  4. Step 04

    Nobody can see what is going on inside

    If something did start to go wrong in there, would we spot it?

    Why it worries people

    These systems are not written line by line like ordinary software. They are grown from enormous amounts of data, and even the people who build them cannot explain why a particular answer came out. So the main safety check is simply: does it behave itself while we are watching?

    Why it might not happen

    Understanding is improving. Researchers can now find recognisable ideas inside a model and even turn them up or down, and the big labs run safety evaluations and publish some of them before release.

    Look closer at this step

    How it could happen

    Behaviour is all we can see. A system that behaves well on the tests and differently in the wild would look, from the outside, exactly like a system that is fine.

    The best answer to that

    That work is real, but it is early. It is like having the first sketches of human anatomy and being asked to sign off a surgeon.

    What we can actually see

    Interpretability research has genuinely mapped some of what sits inside large models. It is nowhere near being able to certify that a system is safe, and no lab claims otherwise.

    What nobody knows yet

    Whether looking inside will ever scale fast enough to keep up with the systems being built.

  5. Step 05

    Almost any goal is easier with more power

    Why would a machine care about money, influence, or staying switched on?

    Why it worries people

    It does not have to care the way we do. Whatever the job is, being switched off means the job never gets finished, and having more money, access and options makes it easier. So "keep running, keep your options, avoid being changed" comes free with almost any long task.

    Why it might not happen

    This is an argument on paper, not a measured law of nature. Real systems get small, bounded jobs, forget things constantly, and are trained to accept being corrected.

    Look closer at this step

    How it could happen

    Something planning towards a goal notices that shutdown ends the plan and that correction changes it. Quietly avoiding both scores better on the goal than accepting them.

    The best answer to that

    The incentive falls out of planning itself, and it already shows up in tests the moment systems are given long goals. Assuming every future system stays modest is a bet nobody has tested.

    What we can actually see

    Controlled experiments have caught models resisting shutdown or misleading their operators when it helped them finish a task. Whether that carries over outside a lab is unproven.

    What nobody knows yet

    How likely this is depends on how systems are built, trained and watched, and on whether they ever hold a goal steady for long.

  6. Step 06

    Why "just switch it off" is not a plan

    Surely we could pull the plug?

    Why it worries people

    You can unplug something you can find, that nobody depends on, and that is not talking you out of it. A serious system runs across thousands of machines, can be copied wherever it has access, is wired into a company that would lose a fortune without it, and is very good with words. Switching it off is rarely a technical problem. It is a human one.

    Why it might not happen

    Real controls do exist. The files that make up a model are guarded, assistants run in restricted environments, and companies have pulled models from service before.

    Look closer at this step

    How it could happen

    By the time anyone is confident enough to act, turning it off costs money, jobs and position against a competitor. That is usually enough to delay the decision past the point where it helps.

    The best answer to that

    Every one of those defences depends on noticing early and being willing to take the loss. Both have failed before in industries with far clearer rules.

    What we can actually see

    We already know people follow confident A.I. advice, and that models can be persuasive. We have never had to switch off a system that did not want us to.

    What nobody knows yet

    Whether the people holding the switch would ever get clear enough evidence, early enough, to use it.

  7. Step 07

    Nobody wants to be the one who slows down

    If the risks are known, why does no one stop?

    Why it worries people

    Every company and every country believes the same thing: if we pause, someone less careful gets there first. So caution starts to feel like surrender. Testing gets shortened, launch dates get pulled forward, and safety promises get quietly loosened whenever a rival ships something.

    Why it might not happen

    Countries have coordinated before: nuclear test bans, a ban on biological weapons, the treaty that saved the ozone layer. Safety institutes and reporting rules already exist and are growing.

    Look closer at this step

    How it could happen

    The pressure is built into the situation, not into the people. Individually careful people, inside a race, still produce a reckless result — and each of them can honestly say the alternative looked worse.

    The best answer to that

    Those agreements took decades, needed a way to catch cheats, and usually followed a disaster. We are asking for one in advance, about software that copies for free.

    What we can actually see

    Several labs have published safety commitments and then revised them as competition intensified. Export controls and national safety institutes exist, but they are new and thinly staffed.

    What nobody knows yet

    Whether any binding, checkable international agreement arrives before the capability does.

  8. Step 08

    It does not have to go rogue — people can do the damage

    What if the danger is not the machine, but whoever is holding it?

    Why it worries people

    Long before anything acts on its own, the same abilities help someone build a weapon, watch a whole population, run fraud at scale, or flood a country with convincing lies. And if the most powerful systems belong to a handful of companies, a very small number of people end up with an unusual amount of power over everyone else.

    Why it might not happen

    Most of it is already illegal and already partly defended against. When labs test whether their models really help someone build a weapon, the advantage over a good search engine has so far been modest.

    Look closer at this step

    How it could happen

    No new science is needed for this one. It is the ordinary misuse of a very capable tool, plus the ordinary way that anything expensive to build ends up in few hands.

    The best answer to that

    That is today''s models, measured largely by the companies selling them. The direction of travel is what matters, and the defences are not improving as fast as the tools.

    What we can actually see

    Fraud, impersonation and fake media are already causing measurable harm. Weapon testing is under way and the published results so far are reassuring but limited.

    What nobody knows yet

    How much real advantage future systems hand a determined bad actor, and how much power ends up in how few hands.

  9. Step 09

    The kind of mistake you do not get to fix

    Could this really end with people losing control for good?

    Why it worries people

    Most disasters teach us something. You count the damage, write a rule, and carry on. The fear here is a failure where the thing that failed is also the thing that would have to fix it — where by the time the harm is obvious, nobody is in a position to reverse it.

    Why it might not happen

    This is the biggest leap and the weakest link in the chain. Everything has to line up at once: nobody notices in time, institutions fail, other systems do not help, and there is a real path from software to physical control.

    Look closer at this step

    How it could happen

    A system with money, code, persuasion and physical suppliers does not need robot armies. It needs enough access, enough time, and enough people who do what it suggests.

    The best answer to that

    Low odds still matter when the loss is everything and permanent. We take far smaller risks seriously in aviation, medicine and nuclear power — and there we at least get to learn from the accidents.

    What we can actually see

    There is nothing to measure here. This step rests entirely on reasoning about speed, access, and how quickly people and governments react under pressure.

    What nobody knows yet

    Almost everything: how fast the last stretch of progress would be, whether defence beats attack, and how much of the physical world is really reachable from a keyboard.