We scan new podcasts and send you the top 5 insights daily.
Claims that AI treaties are unverifiable lack imagination. During the Cold War, the US and USSR agreed to saw bombers in half on runways, allowing spy planes to visually confirm disarmament. Similar 'outside-the-box' physical verification methods, like publicly escrowing or destroying GPUs, could work for AI.
The US and USSR, despite being adversaries, collaborated to prevent nuclear proliferation to rogue actors. A similar model can be applied to AI. The US and China share an interest in preventing powerful open-weight models from being used for cyber-attacks or bio-terrorism by third parties, creating a foundation for a safety dialogue.
To prevent a reckless race, a proposed solution is a U.S.-China treaty to govern the resources needed for frontier AI. This would involve tracking and monitoring advanced AI chips in data centers and imposing a verifiable cap on the computational power used for any single training run.
For a blueprint on AI governance, look to Cold War-era geopolitics, not just tech history. The 1967 UN Outer Space Treaty, which established cooperation between the US and Soviet Union, shows that global compromise on new frontiers is possible even amidst intense rivalry. It provides a model for political, not just technical, solutions.
The path to surviving superintelligence is political: a global pact to halt its development, mirroring Cold War nuclear strategy. Success hinges on all leaders understanding that anyone building it ensures their own personal destruction, removing any incentive to cheat.
A global AI safety regime should learn from nuclear arms control by focusing on the physical infrastructure that enables strategic capabilities. Instead of just seeking promises, it should aim to control access to chokepoints like advanced chip manufacturing and the massive data centers required for frontier models.
The belief that AI development is unstoppable ignores history. Global treaties successfully limited nuclear proliferation, phased out ozone-depleting CFCs, and banned blinding lasers. These precedents prove that coordinated international action can steer powerful technologies away from the worst outcomes.
Instead of supervising an AI's hidden thought process, we can demand it produces a 'certificate of reasoning'—a checkable proof—along with its output. This could include citations or sensitivity analyses, shifting verification from observing the process to checking the provided proof.
International AI treaties, particularly with nations like China, are unlikely to hold based on trust alone. A stable agreement requires a mutually-assured-destruction-style dynamic, meaning the U.S. must develop and signal credible offensive capabilities to deter cheating.
While an AI can deceive humans, it cannot deceive reality. Musk posits that the ultimate reinforcement learning test is to have AI design technologies that must work against the laws of physics. This 'RL against reality' is the most fundamental way to ground AI in truth and combat reward hacking.
International AI treaties are feasible. Just as nuclear arms control monitors uranium and plutonium, AI governance can monitor the choke point for advanced AI: high-end compute chips from companies like NVIDIA. Tracking the global distribution of these chips could verify compliance with development limits.