Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

The effectiveness of sophisticated AI systems doesn't scale linearly with LLM improvements. Instead, they hit a capability threshold. John Platt notes his ERA system was "impossible" with Gemini 2.0 but became "amazing" with version 2.5, demonstrating a phase change in utility.

Related Insights

The dramatic improvements from GPT-2 to GPT-4 were driven by a simple law: bigger models and more training data yielded better results. This trend has stopped. Recent attempts to scale even larger models have produced only marginal gains, forcing the industry into more complex, narrow optimizations instead of giant leaps.

The significant leap in LLMs isn't just better text generation, but their ability to autonomously execute complex, sequential tasks. This 'agentic behavior' allows them to handle multi-step processes like scientific validation workflows, a capability earlier models lacked, moving them beyond single-command execution.

The relationship between computing power and AI model capability is not linear. According to established 'scaling laws,' a tenfold increase in the compute used for training large language models (LLMs) results in roughly a doubling of the model's capabilities, highlighting the immense resources required for incremental progress.

Contrary to expectations of diminishing returns, the pace of AI development is accelerating. Arena's data shows that performance gaps between new model generations (e.g., GPT image 2 vs 1.5) are sometimes the largest in history, suggesting progress is speeding up rather than plateauing.

Google's new state-of-the-art Deep Research agents are still powered by the older Gemini 3.1 Pro model. Their significant performance improvements come entirely from 'harness upgrades' and additional inference techniques. This demonstrates that the systems, tools, and processes surrounding a model are now a primary driver of capability, not just the raw power of the base model itself.

While GPT-5.5 is a massive technical improvement, it may not feel transformative for 99% of users' daily workflows. Previous models like GPT-5.4 were already proficient enough for common tasks. The new model's value is realized at the ceiling of capability, on complex edge-case problems that stressed older models, rather than in everyday use.

Third-party tracker METR observed that model complexity was doubling every seven months. However, a recent proprietary model shattered this trend, demonstrating nearly double the expected capability for independent operation (15 hours vs. an expected 8). This signals that AI advancement is accelerating unpredictably, outpacing prior scaling laws.

The market often misinterprets AI progress as linear. However, a clear 'scaling law' dictates that a tenfold increase in the computing power used to train LLMs results in a twofold capability improvement. This exponential relationship means future advancements will be far more disruptive and surprising than incremental projections suggest.

The true measure of a new AI model's power isn't just improved benchmarks, but a qualitative shift in fluency that makes using previous versions feel "painful." This experiential gap, where the old model suddenly feels worse at everything, is the real indicator of a breakthrough.

While new large language models boast superior performance on technical benchmarks, the practical impact on day-to-day PM productivity is hitting a point of diminishing returns. The leap from one version to the next doesn't unlock significantly new capabilities for common PM workflows.