Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Critics argue that proposed third-party evaluators, such as Meter, lack true independence. Their staff often includes former employees from the very AI labs they would audit (OpenAI, Anthropic), creating a "revolving door" that raises questions about conflicts of interest.

Related Insights

Proposed self-regulatory bodies for AI safety have a built-in flaw: they are incentivized to be overly restrictive. They face all the blame for safety failures but get no credit for economic gains from innovation, leading to a natural bias that stifles progress.

The primary constraint for AI safety organizations like Meter is a lack of technical talent, not access to frontier models. They are in a "state of triage," turning down research opportunities because they lack the staff to pursue critical safety questions, a key vulnerability in the ecosystem.

Despite review from governments and AI labs, the International AI Safety Report's writers were independent contractors for Yoshua Bengio's Mila research institute. This structure ensured they were not obligated to incorporate feedback from powerful stakeholders, preserving the document's scientific integrity.

The AI safety auditing ecosystem is fundamentally flawed. Auditors get minimal access (e.g., three days with Astra), are underfunded, constantly lose talent to the very labs they audit, and face legal and operational hurdles, making effective, independent oversight nearly impossible.

The AI industry has no third-party verification; labs self-report performance on bias and accuracy via blog posts. Campbell Brown likens this to banks auditing themselves, arguing it creates an accountability vacuum and undermines public trust in a foundational technology.

Just as auditors consulting for companies they audited led to scandals like Enron, AI evaluators selling training data to model labs creates a mixed incentive. This encourages a "pay to win the benchmark" culture, which undermines the integrity of the evaluation process and ultimately harms the market.

External investigators into AI incidents, like at OpenAI, face a power imbalance. Their access is limited, and they must stay on good terms with labs to be invited back, compromising the candor of their reports and hindering true oversight.

The AI safety movement is heavily influenced and funded by the Effective Altruism (EA) community, which has deep financial and personal ties to frontier labs like Anthropic. Their push for regulation, framed as a response to existential risk, often masks a commercial incentive to create a protected duopoly.

A safety scorecard reveals that even leading labs like OpenAI and Anthropic are failing at basic, achievable AI control measures. Anthropic, despite its safety-first reputation, notably lacks a clear, pre-written plan for containing a misbehaving AI—a non-technical but critical vulnerability.

Third-party AI safety researchers operate under a significant power imbalance. They have no guaranteed right to access pre-release models and often feel pressured to temper their public criticism to ensure they are "invited back next time," potentially compromising the full transparency of their findings.

"Independent" AI Safety Evaluators Suffer from a Revolving Door with AI Labs | RiffOn