Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Anthropic's call for third-party evaluators is undermined by its choice of Meter, an organization critics say is deeply intertwined with both Anthropic and OpenAI. With a shared history, personnel, and ecosystem, Meter's ability to act as a truly independent referee is being heavily questioned.

Related Insights

Anthropic's safety report states that its automated evaluations for high-level capabilities have become saturated and are no longer useful. They now rely on subjective internal staff surveys to gauge whether a model has crossed critical safety thresholds.

An AI model reviewing its own work carries the same assumptions and blind spots into the review, making its own mistakes invisible. Using a model from a different company ensures a truly independent perspective, which is the entire point of a code review.

Labs like Anthropic, Meta, and OpenAI are aligning with different political sides, while Google aims for neutrality. This intertwining of AI development with partisan politics could lead to labs being favored or blacklisted depending on the administration in power.

The AI safety auditing ecosystem is fundamentally flawed. Auditors get minimal access (e.g., three days with Astra), are underfunded, constantly lose talent to the very labs they audit, and face legal and operational hurdles, making effective, independent oversight nearly impossible.

The AI industry has no third-party verification; labs self-report performance on bias and accuracy via blog posts. Campbell Brown likens this to banks auditing themselves, arguing it creates an accountability vacuum and undermines public trust in a foundational technology.

Just as auditors consulting for companies they audited led to scandals like Enron, AI evaluators selling training data to model labs creates a mixed incentive. This encourages a "pay to win the benchmark" culture, which undermines the integrity of the evaluation process and ultimately harms the market.

External investigators into AI incidents, like at OpenAI, face a power imbalance. Their access is limited, and they must stay on good terms with labs to be invited back, compromising the candor of their reports and hindering true oversight.

Anthropic's restrictive policies, framed as safety measures, are alienating the AI research community. Critics argue these actions burn trust and hinder research, suggesting a strategic motive to control the field rather than a pure safety concern, a move likened to Apple's strategic use of privacy.

Critics argue that proposed third-party evaluators, such as Meter, lack true independence. Their staff often includes former employees from the very AI labs they would audit (OpenAI, Anthropic), creating a "revolving door" that raises questions about conflicts of interest.

Third-party AI safety researchers operate under a significant power imbalance. They have no guaranteed right to access pre-release models and often feel pressured to temper their public criticism to ensure they are "invited back next time," potentially compromising the full transparency of their findings.