Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

An Anthropic alignment lead publicly stated a >10% chance of human extinction from AI within a decade. This creates a paradox: if a company truly believes its work carries such a high risk of global catastrophe, the logical response would be to shut down, not to continue development while trying to solve alignment.

Related Insights

The decision to silently nerf AI research stems from a specific belief in catastrophic risk ("foom"), positioning Anthropic as the gatekeeper of AI progress. This reveals a level of hubris that presumes they can control frontier development without pushback from researchers, enterprises, or governments.

The conversation about AI causing human extinction isn't led by outsiders but by insiders. After a researcher resigned from AI firm Anthropic over safety concerns, his former boss—the head of the AI safety team—publicly agreed, estimating the chance of AI ending the world at a staggering 10%.

Public proclamations of AI-driven extinction, like an Anthropic researcher's 10% odds of human annihilation, may be a deliberate strategy. By presenting worst-case scenarios, these individuals aim to trigger urgent conversations and push the industry and regulators toward implementing stronger safety measures.

While not a consensus, surveys of AI researchers reveal significant concern. The median respondent in a large survey assigned a 5% probability to human extinction or a similar disaster from AI, with a third to a half placing the risk at 10% or higher, suggesting the threat is taken seriously within the field.

Many top AI CEOs openly admit the extinction-level risks of their work, with some estimating a 25% chance. However, they feel powerless to stop the race. If a CEO paused for safety, investors would simply replace them with someone willing to push forward, creating a systemic trap where everyone sees the danger but no one can afford to hit the brakes.

The leaders of top AI labs have signed statements acknowledging AI could cause human extinction. Yet, a safety report gives them failing grades on 'existential safety,' finding it jarring that these same leaders are actively building superintelligence without any articulated plan for how to maintain human control over the technology.

There is a fundamental contradiction when AI leaders publicly warn of existential risks while also aggressively fundraising and preparing for IPOs. If the danger were truly imminent, the logical action would be to halt development, not capitalize on it. This discrepancy suggests financial motives may outweigh safety concerns.

Sam Harris highlights the bizarre cultural phenomenon of AI leaders openly stating high probabilities (e.g., 20%) for existential risk while racing to build the technology. He contrasts this with Manhattan Project scientists, who proceeded only after calculating the risk of igniting the atmosphere as infinitesimal, not a double-digit percentage.

Just before the launch, Anthropic publicly warned about the dangers of frontier AI, urging for a "brake pedal." This cautious positioning, possibly intended to demonstrate responsibility, may have been taken at face value by the government, contributing to the rapid and decisive shutdown of their most advanced model.

Framing AI as an existential threat is a poor marketing strategy for companies like Anthropic. It directly triggers calls from politicians for data center moratoriums, congressional hearings, and bans on superintelligence research—creating immense business risk just before a critical IPO.

Anthropic's >10% AI Extinction Risk Belief Creates a 'Shutdown Paradox' | RiffOn