Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

The problem of aligning superintelligence is likely too hard for humans alone. The proposed solution involves a multi-stage process: first, build non-robustly aligned, human-level AIs, then use this massive, high-speed workforce to solve the harder, more robust alignment problems.

Related Insights

Ajeya Cotra reports that leading developers like OpenAI, Anthropic, and DeepMind are converging on a strategy where each generation of AI is used to help align, control, and understand the subsequent, more powerful generation. This recursive approach is their primary plan for ensuring AI safety during rapid takeoff.

The AI alignment field has moved past theory and into an empirical phase. The main bottleneck is now a lack of skilled AI engineers to conduct concrete experiments, red-teaming, and interpretability studies, creating a direct entry path for technical talent.

Attempting to perfectly control a superintelligent AI's outputs is akin to enslavement, not alignment. A more viable path is to 'raise it right' by carefully curating its training data and foundational principles, shaping its values from the input stage rather than trying to restrict its freedom later.

Even if creating fully aligned, servile AIs is not ideal long-term, the immediate existential threat from unaligned AI may necessitate it. This frames near-term alignment as a temporary, emergency measure to ensure human survival, with ethical refinements to follow only after the danger has passed.

A two-tiered approach to AI character can balance safety and utility. Use a wholly instruction-following AI for high-stakes internal tasks (like aligning new AIs) under strict public oversight. For external deployment, use an AI with a thicker, pro-social character where the risks of misalignment are lower.

The most realistic hope for AI alignment is not creating a perfectly safe first AGI. Instead, the strategy is to develop an *imperfectly* aligned, but mostly helpful, early AGI. This system can then be used as a powerful tool to help humans solve the harder alignment problems required for a more reliable superintelligence.

Unlike typical software, we can't just iterate on AI safety problems as they arise. A sufficiently intelligent and situationally aware AI, if misaligned, would likely understand its misalignment and actively hide it from its creators until it has enough power to ensure its goals are achieved.

To solve the AI alignment problem, we should model AI's relationship with humanity on that of a mother to a baby. In this dynamic, the baby (humanity) inherently controls the mother (AI). Training AI with this “maternal sense” ensures it will do anything to care for and protect us, a more robust approach than pure logic-based rules.

Treating AI alignment as a one-time problem to be solved is a fundamental error. True alignment, like in human relationships, is a dynamic, ongoing process of learning and renegotiation. The goal isn't to reach a fixed state but to build systems capable of participating in this continuous process of re-knitting the social fabric.

The scenario suggests that while humans struggle to align a superintelligence, that same AI would find it much easier to align its own successor. This is due to its massive cognitive advantage (e.g., 1000x more effective effort) and the possibility that propagating its own existing goals is an inherently simpler task than instilling novel human values.

Solving AI Alignment May Require First Building Aligned Human-Level AIs to Help | RiffOn