The sentence was posted on a Tuesday, late, the way people say things they have been carrying for a while. "We really do earnestly believe AI could kill all humans," wrote Evan Hubinger, who leads alignment science at Anthropic. "I personally think it is greater than ten percent within the next decade." The post was seen roughly 9.6 million times. It was not a leak, not a warning buried in a technical appendix, not a philosopher's thought experiment. It was a senior safety researcher at one of the three companies building the technology, saying plainly that his own employer's product carries a one-in-ten chance, or worse, of ending the species.

He was defending a colleague. The day before, Jacob Coxon, a 27-year-old researcher who had spent three years on pretraining work, first at OpenAI and then at Anthropic, announced he was leaving not just the company but the field. His stated reason: "Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives."

Key Points

  • Evan Hubinger, Anthropic's alignment science lead, publicly estimated a greater than 10% chance AI could cause human extinction within a decade.
  • He said risk from current models is low. His concern is future systems capable of recursive self-improvement, for which, he wrote, there is no plan to solve alignment.
  • Jacob Coxon, a 27-year-old pretraining researcher, resigned from Anthropic and left AI entirely, accusing the industry of "gambling with our lives."
  • The exchange followed reports that Anthropic declined, for the first time, to give the UK's AI safety body pre-release access to its latest model.

What the Number Actually Means

Ten percent is not a rhetorical figure. Inside AI safety circles it has a name, p(doom), the probability a given person assigns to catastrophic outcomes from advanced AI. For years these estimates lived in private conversations and long-form forum posts, treated by outsiders as a subculture's morbid parlor game. What changed this week is not the number. Hubinger's estimate sits well within the range that Anthropic's own leadership has floated in interviews. What changed is that a person whose job is to prevent the outcome said the odds out loud, in the first person, and attached them to a timeline.

The distinction Hubinger drew matters. He does not claim the models available today are dangerous in this way. Those, he said, carry low risk. The concern is a specific future capability: systems able to improve themselves, to rewrite and retrain their own successors faster than humans can review the results. For that scenario, he wrote, "we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to." Read carefully, that is an admission that the safety problem the entire industry is premised on solving remains unsolved, and that the people closest to it cannot say when, or whether, it will be.

The Resignation Underneath It

Coxon's departure gave the number a human shape. He was not a policy critic or an outside skeptic. He worked on the core training pipeline. His argument was not that Anthropic's safety research is fake, but that competition renders it insufficient: each lab believes no one else will act responsibly, so each justifies moving fast to reach superintelligence first. The safety work is real. The race around it makes the safety work lose.

"We're on track for a lot of the most aggressive of these scenarios where by the end of next year things could be out of control already," he told The Wall Street Journal. He described the physical reality of the work with something close to disbelief: "It's kind of insane that it has to happen on the MacBooks of some engineers living in San Francisco instead of a bunker in the desert." This is the same anxiety that a Canadian parliamentary committee heard in testimony earlier this year, when researchers warned that loss-of-control scenarios were no longer hypothetical.

Why the Timing Is Not an Accident

The exchange landed days after the Financial Times reported that Anthropic had, for the first time, declined to give the United Kingdom's AI safety institute pre-release access to its newest model for evaluation. For a company that has built its public identity on being the careful lab, the one that would submit to outside scrutiny, that is a meaningful reversal. It is hard to read Hubinger's post and Coxon's resignation as unrelated to it.

The pattern is familiar from earlier reporting. Anthropic's own commissioned work has shown that the evaluation methods used to certify frontier models are already strained past their design limits. When the people who build the tests say the tests are breaking, and the company then withholds a model from the testers, the safety story stops being reassuring and starts being a description of the problem.

The Claim

"We really do earnestly believe AI could kill all humans. I personally think it is greater than ten percent within the next decade." — Evan Hubinger, alignment science lead, Anthropic

The Ethical Weight of Saying It

There is a version of this story that treats Hubinger's post as reckless, a senior figure spreading fear about a technology that, for now, mostly writes code and drafts emails. That reading misses what is actually uncomfortable here. If a researcher genuinely believes the probability is above ten percent, staying quiet would be the unethical choice. The discomfort is not that he said it. The discomfort is that the number is defensible, that his colleagues do not appear to be disputing it, and that the response from the industry has been to keep building at the same pace.

This is the ethical core of the moment. We have moved past the phase where the danger was contested by the builders. The people with the most information now assign double-digit odds to extinction and continue the work anyway, each reasoning that their restraint would only hand the lead to someone less careful. It is a collective-action failure playing out in real time, and no individual inside it has the power to stop it. That includes the ones raising the alarm.

What Is Left to Do

Coxon's exit will not change Anthropic's trajectory, and he seemed to know it. Hubinger's post will not slow the release schedule. The value of both is narrower and still real: they are on-the-record admissions, from inside, that the safety problem is unsolved and the timeline is short. That is the kind of evidence that regulation is eventually built on.

The uncomfortable question for readers is what to do with a ten percent figure. It is too high to ignore and too abstract to act on individually. What it does justify is a lower bar for external oversight, mandatory pre-release evaluation, hard limits on self-improving systems, and treating the labs' own stated probabilities as admissible evidence rather than personal opinion. The people building the technology have told us how the bet is priced. The remaining decision is whether the rest of us accept those odds on their behalf, or insist on being asked first. The same absence of consent runs through nearly every story on this site, from who gets to redesign human relationships to who gets to wager the species.