Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Analyst Gavin Baker suggests that embedding third-party evaluators is a savvy legal move for AI companies. It demonstrates a "duty of care," which can help limit liability in future lawsuits over model outputs, much like Section 230 protected early internet companies.

Related Insights

The emergence of powerful, uncensored open-weight models like Obliteration.ai's demonstrates that safety guardrails from companies like OpenAI are easily bypassed. This suggests the long-term solution for AI safety won't be technical restrictions at the model level, but rather legal and regulatory enforcement.

Implementing AI safety guardrails is not cost-prohibitive. The most impactful step, having a second AI model review the primary agent's work, is also the cheapest, accounting for only about 3% of total API costs in the author's experience. This makes it the most efficient first step for improving reliability.

The proper division of labor in AI safety is for the government to define what it's afraid of—the "rules" against biohacking or cyber hacking—and enforce them. The government is not equipped to perform the complex, fast-moving technical work of evaluating if models can break those rules, which should be handled by specialized third-party evaluators.

Demis Hassabis argues that market forces will drive AI safety. As enterprises adopt AI agents, their demand for reliability and safety guardrails will commercially penalize 'cowboy operations' that cannot guarantee responsible behavior. This will naturally favor more thoughtful and rigorous AI labs.

Illinois's new AI safety law introduces a key accountability measure missing from other state regulations: required independent, third-party audits of major AI systems. This move, supported by OpenAI and Anthropic, establishes a stronger framework for external oversight of AI safety.

An FDA-style regulatory model would force AI companies to make a quantitative safety case for their models before deployment. This shifts the burden of proof from regulators to creators, creating powerful financial incentives for labs to invest heavily in safety research, much like pharmaceutical companies invest in clinical trials.

Critics argue that proposed third-party evaluators, such as Meter, lack true independence. Their staff often includes former employees from the very AI labs they would audit (OpenAI, Anthropic), creating a "revolving door" that raises questions about conflicts of interest.

The U.S. has a built-in mechanism for AI safety that precedes formal regulation: the court system. The potential for lawsuits (tort law) incentivizes model makers to act responsibly, acting as a form of self-regulation that doesn't require a slow-moving government bureaucracy.

A straightforward regulatory step would be to hold AI companies legally responsible for any crimes their models commit. This simple shift in liability would force labs to slow down and prioritize safety, as they would be unwilling to deploy models they cannot fully control.

Responsibility for medical AI safety is dangerously diffuse. Foundation model creators do basic checks, and application builders make their own claims, but no independent party verifies performance. This "everyone is responsible, so no one is responsible" paradox leaves patients and hospitals vulnerable, creating a critical need for a neutral, third-party referee.