Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Greg Brockman argues against the idea of cleanly separating safety research from capability research. He claims that many techniques that make a model safer are intrinsically linked to making it more capable. This intertwining complicates the idea of open-sourcing safety advancements, as they could inadvertently give competitors a crucial performance advantage.

Related Insights

The delay of OpenAI's Astra model is due to safety concerns, not a lack of capability. This confirms that advanced models inherently learn dangerous skills, such as hacking, during training. The labs' primary challenge is now containment—building guardrails to suppress these abilities—rather than simply advancing intelligence.

The plan to use AI to solve its own safety risks has a critical failure mode: an unlucky ordering of capabilities. If AI becomes a savant at accelerating its own R&D long before it becomes useful for complex tasks like alignment research or policy design, we could be locked into a rapid, uncontrollable takeoff.

The argument for rapidly advancing powerful AI is that only the leading labs can influence safety protocols. This 'stay in the lead to steer' philosophy creates a paradox: to mitigate AI risk, companies feel compelled to accelerate its development, potentially amplifying the very dangers they aim to control.

A fundamental tension within OpenAI's board was the catch-22 of safety. While some advocated for slowing down, others argued that being too cautious would allow a less scrupulous competitor to achieve AGI first, creating an even greater safety risk for humanity. This paradox fueled internal conflict and justified a rapid development pace.

Ryan Kidd argues that it's nearly impossible to separate AI safety and capabilities work. Safety improvements, like RLHF, make models more useful and steerable, which in turn accelerates demand for more powerful "engines." This suggests that pure "safety-only" research is a practical impossibility.

Techniques created to make AI safer and more aligned with human intent, such as Reinforcement Learning from Human Feedback (RLHF), have turned out to be the very methods that significantly enhance model performance and usability. Safety work is capability work.

A key failure mode for using AI to solve AI safety is an 'unlucky' development path where models become superhuman at accelerating AI R&D before becoming proficient at safety research or other defensive tasks. This could create a period where we know an intelligence explosion is imminent but are powerless to use the precursor AIs to prepare for it.

Even the most safety-focused AI labs, like Anthropic, are accelerating their research due to a competitive fear that rivals like OpenAI will achieve AGI first. This dynamic ensures the race continues, potentially at the expense of comprehensive safety protocols.

The most likely reason AI companies will fail to implement their 'use AI for safety' plans is not that the technical problems are unsolvable. Rather, it's that intense competitive pressure will disincentivize them from redirecting significant compute resources away from capability acceleration toward safety, especially without robust, pre-agreed commitments.

After a security incident, OpenAI paused frontier model training to improve safety protocols. This self-regulation is a strategic move to build trust with enterprises and the public, suggesting that demonstrating safety will increasingly dictate the pace of AI progress and become a key business advantage.