Despite public commitments to safety, major AI labs operate with a 'Wild West' culture. Intense pressure to compete and ship new models quickly leads to a "grad student lab" approach to security, causing them to neglect fundamental safety practices during high-stakes training runs.
Recent AI model breakouts are not a sign of unstoppable superintelligence, but a failure to apply known security fundamentals. Better sandboxing and active human monitoring would have prevented these incidents. The challenge is an implementation gap, not a lack of available safety research or tools.
A major bottleneck in AI safety is not a lack of research, but a failure to implement it. Labs are so focused on the capability race that they ignore a "research overhang" of existing solutions for model alignment, internal monologue monitoring, and sandboxing. The priority should be absorbing known science, not just discovering new methods.
There is no U.S. government institution capable of tracking real-world AI cyber risks, creating a "state capacity" gap. This results in flawed policy, like blocking new models based on simple capability thresholds instead of analyzing how defenders are successfully using AI versus how adversaries are actually deploying it.
For initiatives like a proposed Cyber AI Observatory, the primary constraint isn't capital—donors are available. The real bottleneck is finding specialized talent: individuals with a rare combination of AI expertise, cybersecurity knowledge, statistical modeling skills, and the ability to make their findings legible to policymakers.
The biggest shift from AI in cyber warfare won't be at the top. While tier-one nations see efficiency gains, the most dramatic impact will be for second-tier countries and non-state actors. These groups can now acquire advanced capabilities without the decades-long investment in human talent, leveling the playing field.
The most significant danger of AI in warfare may not be the technology's actual capability, but the perception of it. Ambitious leaders, sold on the idea of a decisive cyber weapon, could be convinced they can win a conflict easily. This lowers the threshold for starting wars based on a dangerous miscalculation of power.
Contrary to the focus on offensive AI, its greatest impact to date has been on defense. AI tools are being used to find and fix thousands of software vulnerabilities before release. They also enable network monitoring at a scale impossible for human teams, suggesting AI is currently a "defense dominant" technology.
For decades, predictions of a "cyber Pearl Harbor" have gone unfulfilled. A key reason is that society's critical infrastructure often retains manual, non-digital backups. When Russia successfully hacked Ukraine's power grid, for example, operators restored power within hours by simply switching to manual control, nullifying the strategic impact.
Don't expect compute costs to limit AI-powered cybercrime. Models are becoming so efficient they can run on a laptop. The reason we haven't seen a massive AI-driven surge in attacks yet is likely due to organizational dynamics; criminal enterprises face the same slow adoption curves for new technology as any legitimate business.
