AI models can solve complex, benchmarkable problems like advanced math or chess, yet their overall real-world impact remains limited. This suggests a persistent gap between specialized capabilities and true, world-altering generalization, a modern version of Moravec's paradox where hard problems are easy and easy problems are hard.
Progress in AI isn't a smooth, continuous line. Just as Moore's Law required discrete inventions, AI scaling relies on paradigm shifts. The current Transformer+RL approach may hit diminishing returns, and an AI trained within this paradigm is unlikely to discover the next fundamental breakthrough required to maintain progress.
Even if AIs can automate all technical aspects of research and engineering, humans will retain the crucial role of defining what the AI should actually do. Specifying goals, shaping behavior (like helpfulness), and setting constitutions will be the longest-lasting human contribution to AI development.
The ability to distill the capabilities of a frontier AI model into a smaller, cheaper one is the main factor preventing an oligopoly. Even if a lab's model is continuously improving, competitors can continually distill its public-facing behaviors, ensuring that no single company can maintain an insurmountable lead.
Simply having API access to a frontier model is insufficient for effective distillation. The real competitive advantage lies in possessing a wide, diverse, and realistic distribution of user prompts. This data reveals the model's capabilities on real-world tasks, making it the most valuable asset for training a competitive model.
It is relatively easy to create difficult, puzzle-like environments for AI training. The much harder task is to simulate realism—the complex, multi-turn, multi-objective nature of real-world interactions. Frontier labs that succeed are those that push heavily on the realism axis, which is harder to replicate than benchmark performance.
The path to recursive self-improvement won't start with an AI discovering principles from scratch. Instead, it will begin by automating the process of incorporating the latest human-driven progress. AI labs will turn their recent bug fixes and discoveries into new training environments, effectively distilling the last few months of human R&D into the next model.
Current AI models require thousands of interactions to learn a new skill, making direct learning from real-time human feedback impractical. This inefficiency forces labs to simulate tasks and human interactions within a data center to generate the necessary volume of training data. As sample efficiency improves, learning from live deployment will become more viable.
While models don't learn in real-time from every user, a slower feedback loop is already in place. Labs collect data from deployed models and incorporate it into the pre-training and mid-training of subsequent generations. This constitutes a batched, asynchronous version of a continually learning 'hive mind' intelligence.
AI research is a 'cumulative' task where discoveries (like a new architecture) can be added to a stack and retained. In contrast, many real-world tasks (e.g., being a legal associate) are 'non-stationary,' requiring constant adaptation to changing social dynamics. This could paradoxically make automating AI research easier than creating a truly adaptive real-world agent.
Current methods for updating models with new data suffer from catastrophic forgetting, forcing labs to periodically retrain models from scratch. This inability to continuously integrate new information into the same model is a fundamental technical and economic bottleneck that prevents the creation of a single, persistently learning AI.
A study that trained old models on new datasets (and vice versa) found that data improvements accounted for a 9x gain in compute efficiency, while architectural changes only provided a 3x gain. This suggests that data quality, filtering, and synthesis have been the predominant drivers of progress in pre-training, at least at smaller scales.
Reinforcement Learning is effective because it applies a few, high-signal bits of feedback ('correct' or 'incorrect') to an already capable model. Unlike Supervised Fine-Tuning (SFT), which is noisy and tries to match every token, RL isolates the most crucial learning signal. This allows for efficient tweaking of a model's policy without reteaching it everything.
While RL may not provide perfect cross-domain reasoning (e.g., math to code), its key contribution is teaching models 'horizon generalization.' This is the ability to use more tokens productively over a longer period to make progress on a complex task. This meta-skill is a primary driver of recent capability improvements and appears to be doubling every three months.
