Eddy Lazzarin argues that today's AI incidents, like hacking, are not early signs of rogue superintelligence. Instead, they are familiar cybersecurity and control failures that can be addressed with existing tools like cryptography and better system design, rather than abstract "alignment" work.
To counter AI that might fake alignment, treat it like a human liar. Engage in "iterated games" to test its behavior over time. This allows for developing reputation systems for specific models, enabling users to distinguish trustworthy AI from unreliable ones, much like we do with people and institutions.
The concept of independent AI evaluators seems positive but harbors a hidden risk. If these evaluators all come from the same social and ideological circles, the system becomes merely "distributed," not truly "decentralized." This can lead to a single cultural unit subtly controlling a critical industry under the guise of safety.
The current AI debate is lopsided, focusing heavily on the "probability of doom" (P-doom). Eddy Lazzarin advocates for equally considering the "probability of abundance" (P-abundance)—the immense potential benefits we risk losing. The opportunity cost of delaying progress must be a central part of the safety conversation.
The argument against pausing AI development is that capability improvements are a prerequisite for better safety, not an obstacle. More advanced models will provide the tools needed for superior mechanical interpretability ("mech interp"), meaning progress itself is the path to safer systems.
Pursuing a future with zero malicious AI is a futile goal. Bad actors will inevitably create unaligned models. The only realistic and effective countermeasure is to accelerate the development of powerful "good" models, which will be necessary to defend against and control the bad ones, similar to how cybersecurity operates.
The intensity and strangeness of the current AI debate stem from niche ideas—long confined to Silicon Valley subcultures like Effective Altruism—suddenly entering the mainstream political arena. This "worlds colliding" moment is forcing these abstract arguments to be reformulated as they are exposed to broader public scrutiny and reality.
