We scan new podcasts and send you the top 5 insights daily.
Counter-intuitively, new models like GPT-6-1-Soul show degraded performance on the highest effort settings. This suggests frontier models are hitting a complexity ceiling where more processing leads to 'overthinking' and second-guessing correct answers, forcing new optimization approaches beyond just scaling up.
The dramatic improvements from GPT-2 to GPT-4 were driven by a simple law: bigger models and more training data yielded better results. This trend has stopped. Recent attempts to scale even larger models have produced only marginal gains, forcing the industry into more complex, narrow optimizations instead of giant leaps.
The enormous compute budget for the original AlphaGo was not about finding the most efficient training method, but about proving a method could work at all. Once a breakthrough is made and the path is clear, subsequent efforts can focus on optimization and achieve similar results with far less compute.
The original playbook of simply scaling parameters and data is now obsolete. Top AI labs have pivoted to heavily designed post-training pipelines, retrieval, tool use, and agent training, acknowledging that raw scaling is insufficient to solve real-world problems.
Over two-thirds of reasoning models' performance gains came from massively increasing their 'thinking time' (inference scaling). This was a one-time jump from a zero baseline. Further gains are prohibitively expensive due to compute limitations, meaning this is not a repeatable source of progress.
Like human experts, advanced AI models improve their answers the more time they spend on a problem. This 'inference scaling' means short evaluations may fail to capture a model's true capabilities, as performance continues to increase with more computation, making it difficult to establish a performance ceiling.
The "effort" setting is not a control for processing time. Instead, it is an input that prompts the model to follow a pre-trained behavior. High effort causes the model to generate more reasoning tokens and tool calls, making it more thorough and certain before it considers a task complete. This behavior is baked into its frozen weights.
The rapid, step-change improvements in LLMs are likely slowing down. This is because models have already been trained on most of the available internet, and the compute budget required for each incremental improvement is increasing exponentially to an unsustainable degree. A new architectural breakthrough, not just more data and compute, is needed for the next leap.
The dominant AI strategy of building increasingly larger models is becoming unsustainable. The primary constraint is memory, which is described as "already broken." Consequently, leading companies are abandoning the "scaling hypothesis" and shifting focus to more efficient models, a paradigm shift from the brute-force approach of the last five years.
Pushing models like Anthropic's Opus 5 to 'max effort' settings can backfire. Performance on some benchmarks peaked at lower settings, as maximum effort can lead to overthinking, unnecessary changes, or 'endless self-verification loops.' This suggests that more compute doesn't always equal better results, requiring users to tune effort levels for optimal outcomes.
Counterintuitively, setting GPT-6 Astra to 'max effort' for some benchmarks led to worse performance. This suggests a novel failure mode for advanced AI, where excessive computation can cause the model to overcomplicate simple tasks, get sidetracked, and deliver suboptimal results, highlighting that more 'effort' is not always better.