Benchmarks become obsolete when models master them, but there's a subtler reason for deprecation: the world changes. For a legal benchmark to remain relevant, it must incorporate updated case law, just as a doctor must stay current with medical knowledge. Benchmarks need to be dynamic snapshots of reality.
Choosing a mid-tier model to save money can backfire. Anthropic's 'cheaper' Sonnet model is sometimes more expensive than the premium Opus model because it is more token-hungry for certain tasks. Cost-per-token is a poor proxy for total cost; performance-based evaluation is essential for determining true ROI.
The proper division of labor in AI safety is for the government to define what it's afraid of—the "rules" against biohacking or cyber hacking—and enforce them. The government is not equipped to perform the complex, fast-moving technical work of evaluating if models can break those rules, which should be handled by specialized third-party evaluators.
As sovereign AI initiatives grow, the risk of an unchecked capabilities race increases. A shared language of evaluations could provide a "trust but verify" mechanism, akin to nuclear arms verification treaties. This allows nations to audit each other's AI for risks like runaway RSI, creating a basis for international policy.
The ideal test for recursive self-improvement (RSI)—letting a model train its next version—is currently infeasible due to cost and time. Instead, VALS measures RSI by creating proxy benchmarks for discrete steps in the model development process, such as pre-training and post-training research, to gauge a model's capabilities.
When a Fortune 10 company gave engineers a $100/day AI tool budget, it created a "dead period" in the afternoon once limits were hit. The most productive hours became 4-6 p.m., when rate limits reset. This anecdote illustrates how companies are mis-valuing AI intelligence, leading to distorted work patterns.
When Meta released Llama 4, it excelled on public benchmarks where questions were open source. However, on VALS' private, held-out benchmarks, it significantly underperformed, revealing a major disconnect and showing that self-reported, public scores can be misleading indicators of true capability.
Early benchmarks like ImageNet used millions of simple input-output pairs. Evaluating modern, agentic AI for tasks like coding an entire application requires the opposite: a small number of test cases (e.g., 50 apps) evaluated against a deeply complex rubric, a trend that will continue as AI tackles more complex work.
Just as auditors consulting for companies they audited led to scandals like Enron, AI evaluators selling training data to model labs creates a mixed incentive. This encourages a "pay to win the benchmark" culture, which undermines the integrity of the evaluation process and ultimately harms the market.
