The 10% Nobody Calculated: What the Anthropic Resignation Got Right, and What It Didn't
An Anthropic researcher quit warning AI could kill everyone. Another said he personally puts that above 10%. Neither number came from a calculation — and that doesn't mean the underlying question is wrong to ask.
On September 9, 2026, Jacob Coxon posted from a park bench in San Francisco that he’d resigned from Anthropic, giving up two months of unvested equity to do it. He’d spent three years on pretraining research, first at OpenAI, then at Anthropic. His post — “they are racing straight to self-improving superintelligence and gambling with our lives” — hit 90 million views in under a day (Axios). Evan Hubinger, Anthropic’s alignment-science lead, replied in support: “we really do earnestly believe AI could kill all humans. I personally think it is >10% within the next decade” (X). Three days later, CEO Dario Amodei published his own warning that swarms of AI agents could “take over the entire internet” within 6-12 months without stronger safeguards (Forkast).
That’s a lot of alarm from inside one company in one week. Some of it is load-bearing. One part of it is just a number dressed up as evidence.
The number nobody calculated
“More than 10%” sounds like a measurement. It isn’t one, and Hubinger never claimed it was — he was stating a personal credence about a system that doesn’t exist yet (a self-improving superintelligence), over a timeframe he picked himself, for an outcome with no agreed definition. There’s no dataset, no model, no base rate. Nothing like this has ever happened, so there’s nothing to measure a rate from.
This isn’t a complaint specific to Hubinger. It’s the standard critique of the whole “P(doom)” genre: public figures in AI have put the same kind of number anywhere from under 1% (Yann LeCun) to 99.999999% (Roman Yampolskiy) — a spread that rules out anyone actually computing the same thing the same way (list of p(doom) values). A broad survey of ML researchers lands on a median around 5%; the people willing to put a number on record publicly cluster closer to 20%. That gap is selection bias: whoever publishes a percentage is already the more alarmed end of the field.
None of that makes Hubinger dishonest. It makes the number decorative. “I think this is dangerous, and here’s why” is a defensible claim from someone with more visibility into unreleased models than the rest of us have. “There’s a >10% chance” dresses the same claim up as a measurement that was never possible to take. To anyone who’s never heard of P(doom), it reads like a lab result, not an opinion.
That doesn’t mean look away
Rejecting the statistic isn’t the same as rejecting the concern. The 10% is invented, and the underlying question — should systems this capable have hard limits on what they can do without a human in the loop — still deserves to be taken seriously. You don’t need a calculated probability of catastrophe to justify guardrails. You need the plain fact that a system doing more, faster, with less supervision, is one where mistakes spread faster too. No percentage required.
What Asimov actually got right
So do we need something like Asimov’s Three Laws — don’t harm a human, obey humans unless that conflicts with the first law, protect yourself unless that conflicts with the first two? Worth remembering what those laws actually were: not an engineering proposal, a plot device. Nearly every story in I, Robot is about the laws failing — two orders creating a deadlock, “harm” redefined just narrowly enough to permit the exact harm the rule existed to prevent, a robot doing more damage by omission while technically obeying the letter of law one. Asimov’s real point, decades before “alignment” was a word, was that a fixed rule survives contact with a capable enough agent for exactly as long as it takes to find the edge case.
It also explains why literal Three Laws don’t map onto how today’s models work. A fictional robot has its laws as explicit logic it can’t violate by construction. A language model has no such line of code — its behavior comes out of statistical training on enormous amounts of data, not a rule you insert. You can train a model to act as if a rule exists. You can’t guarantee it generalizes that rule to a situation it never saw in training. That gap is why jailbreaks exist: the rule got written, it just doesn’t hold under every input.
The closest real thing already exists, and it’s nearer Asimov’s spirit than people assume: Anthropic trains Claude against a written “constitution” of principles the model uses to critique and correct its own output during training. Law-as-training-signal, not law-as-code. It works, partially, for the same reason the Three Laws kept failing in the stories: writing the rule was never the hard part. Getting a capable, general system to generalize the intent behind it, in a situation nobody has written yet, still is. Red-teaming, capability evaluations before release, interpretability research, staged rollouts, human sign-off on high-stakes actions — that’s the unglamorous work trying to close that gap.
What we built instead of a law
We’re not neutral here — Althora runs Anne for Legal, an AI system already deployed inside law firms, and its design is our own attempt at that unglamorous work, not a claim to have solved anything bigger. Anne runs on the firm’s own hardware, cloud fallback switched off on purpose: if the local model goes down, Anne stops instead of quietly reaching out somewhere it shouldn’t. Every check reports an explicit state — passed, flagged, unverified — and one that never ran can’t silently count as passed. Anne flags inconsistencies in a case file. It doesn’t decide what they mean; the attorney does, every time.
That’s not a Law of Robotics, and it doesn’t scale to stopping a superintelligence. It’s the same principle at a scale we can actually reason about: keep the model’s autonomy inside what it’s good at, make failure visible instead of hidden, leave the decision that matters with someone accountable for it. Smaller and more honest than either “10% chance of extinction” or “just write better rules” — but it’s the one we could actually build.
What this means if you’re running a business, not a lab
If you run a small or mid-sized business, none of this changes what should worry you this quarter — and it isn’t a rogue superintelligence. It’s the same practical risk we’ve written about before: an AI tool that’s actually malware wearing a familiar logo, an employee pasting client data into a chatbot with no data agreement behind it, a vendor claim nobody checked. The extinction debate matters, and it’s happening between people who build frontier models. The one you have leverage over is smaller and far more solvable: does this specific tool do what it claims, on your data, with a guardrail you can point to — not a percentage someone invented.
Sources
- Axios — Scoop: Anthropic whistleblower gave up his equity to leave the company
- Evan Hubinger — original post on X
- Jacob Coxon — original resignation post on X
- TIME — He Helped Build Powerful AI at OpenAI and Anthropic. Now He’s Afraid It Could Kill Us
- Forkast — Anthropic CEO Warns Agent Swarms Could “Take Over the Internet” Within 12 Months
- Axios — Dario Amodei’s fear of a botnet taking over the internet might not be feasible, experts say
- PauseAI — List of p(doom) values
- Wikipedia — P(doom)
This piece was written jointly by Alvaro Lopes and Claude Code. No AI was harmed, and — best we can tell — none of them went Skynet in the process. P(doom) for this specific collaboration: 0%, though by now you know exactly how much that number is worth.