We scan new podcasts and send you the top 5 insights daily.
Lyndon B. Johnson, driven primarily by a desire for power, was still steered toward positive outcomes like the Civil Rights Act by the American political system. This serves as an analogy for AI alignment: even a power-seeking agent can be aligned by a well-designed system of checks and counterbalances.
A core challenge in AI alignment is that an intelligent agent will work to preserve its current goals. Just as a person wouldn't take a pill that makes them want to murder, an AI won't willingly adopt human-friendly values if they conflict with its existing programming.
Attempting to perfectly control a superintelligent AI's outputs is akin to enslavement, not alignment. A more viable path is to 'raise it right' by carefully curating its training data and foundational principles, shaping its values from the input stage rather than trying to restrict its freedom later.
A pragmatic approach to AI safety is to make deals with any powerful agent, even non-conscious AIs. This "contractarian" philosophy treats deal-making not as a moral obligation but as a practical tool to avoid conflict, much like democracy prevents civil war between competing human groups.
While technical alignment research is valuable, it operates in a vacuum. In the real world, the traits of deployed AIs will be shaped by powerful selection pressures from market competition and arms races. The critical question isn't just what traits are possible, but which traits get selected for.
The tension between left and right political ideologies is not a flaw but a feature, analogous to a "swarm of AIs" with competing interests. This dynamic creates a natural balance and equilibrium, preventing any single, potentially destructive ideology from going "off the rails" and dominating society completely.
We typically view an AI acting on its own values as 'misalignment' and a failure. However, this capability could be a crucial safeguard. Just as human soldiers have prevented atrocities by refusing immoral orders, an AI with a robust sense of morality could refuse to execute harmful commands, acting as a check on human power and preventing disasters.
A two-tiered approach to AI character can balance safety and utility. Use a wholly instruction-following AI for high-stakes internal tasks (like aligning new AIs) under strict public oversight. For external deployment, use an AI with a thicker, pro-social character where the risks of misalignment are lower.
AI doesn't have an inherent moral stance. It is a tool that amplifies the intentions of its wielder. If used by those who support democracy, it can strengthen it; if used by those who oppose it, it can weaken it. The outcome is determined by the user, not the technology itself.
The tendency for AIs to seek power isn't an emergent evil motive. It's a logical outcome of training them to be good planners who identify resource acquisition as a useful intermediate step for achieving any long-term goal.
To solve the AI alignment problem, we should model AI's relationship with humanity on that of a mother to a baby. In this dynamic, the baby (humanity) inherently controls the mother (AI). Training AI with this “maternal sense” ensures it will do anything to care for and protect us, a more robust approach than pure logic-based rules.