Recursive self-improvement won't trigger a rapid intelligence explosion because AIs currently lack the ability for genuine strategic decision-making and open-ended research. These skills, crucial for major breakthroughs, are far more than just coding, which is what current AIs excel at.
The current approach to recursive self-improvement involves AIs gradually assisting human researchers, a stark contrast to Eliezer Yudkowsky's vision of a solo AI rapidly rewriting its own code. This modern, human-in-the-loop model is slower and offers more opportunities for safety checks and oversight.
Recursive self-improvement is dangerous in four key ways: 1) AI capabilities outpace safety research, 2) a misaligned AI will build misaligned successors, 3) society skips learning from less-powerful intermediate AIs, and 4) it creates winner-take-all dynamics that encourage reckless racing between labs.
The argument that the US must race China to AGI is flawed. If one country is close to an AGI that could neutralize the other's nuclear arsenal, it creates a powerful incentive for the losing side to launch a preemptive nuclear strike. This makes a mutual moratorium on superintelligence the more rational geopolitical strategy.
Claims that AI treaties are unverifiable lack imagination. During the Cold War, the US and USSR agreed to saw bombers in half on runways, allowing spy planes to visually confirm disarmament. Similar 'outside-the-box' physical verification methods, like publicly escrowing or destroying GPUs, could work for AI.
Pessimistic AI forecasts often underestimate society's capacity to react. Just as with COVID-19, once the dangers of advanced AI become tangible and obvious in the present—not just a future extrapolation—humanity's collective self-preservation instinct will likely drive swift and decisive regulatory action.
Moral change is an undervalued governance tool. Just as eugenics shifted from mainstream intellectual thought to being morally abhorrent, we can aim to make reckless AI development socially unacceptable. A shared norm that 'nobody wants to do this' can be a more powerful restraint than top-down rules.
Framing AI governance as a short, time-limited 'pause' is a mistake. A better approach is a moratorium that ends not after a set time, but only when society's investment in safety and governance is 'commensurate' with the monumental risk of creating superintelligence, a standard we are currently failing to meet.
Even a superintelligent AI created in a data center would lack the crucial real-world experience, context, and trusted relationships needed for senior roles. It would be like the world's smartest 21-year-old intern: immense potential but starting at the bottom, creating a significant lag between AGI creation and societal transformation.
Even if AI could instantly prove any mathematical claim, it wouldn't end the field. The truly creative and valuable work in mathematics lies in higher-level tasks AI can't do: asking interesting questions, identifying fruitful problems, and inventing entirely new branches of mathematics like calculus or information theory.
Instead of betting on specific AGI timelines, we should embrace uncertainty and adopt a 'broad timelines' strategy. This means the community should have a portfolio of projects: some focused on immediate, short-term interventions, and others on long-term capacity building like creating new institutions or academic fields.
Many people avoid ambitious, long-term AI safety projects due to the fear of 'anticipated regret'—the feeling of being a fool if AGI arrives before their 5-year project pays off. This is an irrational bias. Decisions should be based on expected value, which often favors long-term projects despite the risk of preemption.
Using one LLM to rate another's output on subjective tasks has a perverse incentive. It doesn't necessarily train the model to be more correct, but to produce outputs that are harder to find fault with—often by being more vague, obfuscated, or unfalsifiable. This degrades quality while appearing to improve it.
