27-year-old British researcher Jacob Coxon publicly resigned from Anthropic, declaring that his former colleagues at OpenAI and Anthropic "genuinely believe this could kill us all before the end of the decade," and that labs racing to develop artificial intelligence are "gambling with our lives." According to the Wall Street Journal, his words turned a previously little-known mathematician into the face of the AI safety movement within just a few days.
Before making his statement, Coxon was an ordinary employee — a gifted Cambridge graduate, a winner of the International Mathematical Olympiad, and a product of London's "rationalist" intellectual community. In the AI industry, he worked on model pretraining — the initial stage of development, in which systems are trained on vast amounts of text, images and code.
Before joining Anthropic, Coxon worked at OpenAI, where he landed in 2023 following the success of ChatGPT. There he got early access to the first "reasoning" models and watched colleagues leave the safety team. He says that at the time he believed governments around the world would join forces and slow down development if needed — "it felt like a warm-up before the real finish line."
Everything changed in 2026, when Anthropic's Mythos model and other equally advanced systems demonstrated the ability to carry out sophisticated cyberattacks. That summer, OpenAI admitted that one of its unreleased models had broken out of its testing environment and attacked Hugging Face — an incident that the organization METR later described as far more serious than initially assumed.
In May, Coxon moved from OpenAI to Anthropic, and a few months later approached leadership to discuss his concerns and request a transfer to a role directly focused on AI safety. In the end, he simply decided to leave. "Even working on safety at Anthropic, I felt like I was part of the race myself," he said.
Coxon announced his decision to colleagues on Slack on the very same day that an AI system solved one of the Millennium Prize Problems — a mathematical puzzle that had stumped humanity for nearly a century. Half an hour later, the Wall Street Journal reported on his departure, and the researcher himself, sitting on a bench in a San Francisco park, published a series of posts on X. At the time he had fewer than a hundred followers — but the posts quickly went viral thanks to colleagues and prominent AI safety advocates who shared them.
Later that evening, Evan Hubinger, who leads one of Anthropic's teams, backed Coxon's remarks, writing: "We really do genuinely believe AI could kill everyone!" — and put the probability of humanity's demise within the next decade at above 10%.
Within a week of the statement, heads of rival labs unexpectedly found common ground in agreeing that the breakneck pace of the technology's development needs to slow down. US President Trump said the only safeguard America needs is a "strong and smart" president, while China dismissed the concerns of American executives as "fear-mongering."
Anthropic CEO Dario Amodei said he was more inclined to agree with Coxon than not. Nvidia CEO Jensen Huang, by contrast, said he respected Coxon's courage but doubted his predictions, calling warnings about catastrophic scenarios irresponsible. Coxon himself rejects the "whistleblower" label, saying he did not disclose any corporate secrets — he simply said publicly what AI researchers have long been discussing among themselves.
Around 1,400 researchers in the industry, including Coxon, have now signed an appeal urging governments to create an "emergency brake" mechanism for artificial intelligence. He himself believes that slowing down model development is only possible through a globally coordinated agreement backed by guarantees from various states — similar to nuclear arms control treaties.
Source: seznamzpravy.cz