Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

AI model improvements have shifted from revolutionary leaps (e.g., high-school to college-level intelligence) to marginal gains. The current difference between top models is akin to a PhD student getting an 'A' versus a 'B'—an improvement that is irrelevant for the majority of everyday tasks that a 'college kid' model can handle perfectly well.

Related Insights

On financial analyst benchmarks, top models from Anthropic, Google, and OpenAI are now almost indistinguishable in capability. This convergence suggests the frontier is commoditizing, questioning the return on investment for massive training runs and shifting value up the application stack.

The dramatic improvements from GPT-2 to GPT-4 were driven by a simple law: bigger models and more training data yielded better results. This trend has stopped. Recent attempts to scale even larger models have produced only marginal gains, forcing the industry into more complex, narrow optimizations instead of giant leaps.

Despite perceptions of rapid acceleration, a large-scale analysis by Google DeepMind and EPOC that stitches together many benchmarks over time shows that general AI capability progress has been remarkably linear. This suggests AI is currently a better tool, not an expanding population of researchers.

While AI progress is marketed in revolutionary "step-changes" (e.g., GPT-3 to GPT-4), the underlying reality is more like compounding interest. A continuous stream of small, incremental improvements are accumulating, and their combined effect is what creates the feeling of an exponential leap in capability over time.

Current AI models resemble a student who grinds 10,000 hours on a narrow task. They achieve superhuman performance on benchmarks but lack the broad, adaptable intelligence of someone with less specific training but better general reasoning. This explains the gap between eval scores and real-world utility.

Broad improvements in AI's general reasoning are plateauing due to data saturation. The next major phase is vertical specialization. We will see an "explosion" of different models becoming superhuman in highly specific domains like chemistry or physics, rather than one model getting slightly better at everything.

For sophisticated users, the performance difference between top AI models is negligible. They are all "so good" that the limiting factor for value creation is no longer the tool's capability, but the user's creativity, time, and attention.

While GPT-5.5 is a massive technical improvement, it may not feel transformative for 99% of users' daily workflows. Previous models like GPT-5.4 were already proficient enough for common tasks. The new model's value is realized at the ceiling of capability, on complex edge-case problems that stressed older models, rather than in everyday use.

The perceived plateau in AI model performance is specific to consumer applications, where GPT-4 level reasoning is sufficient. The real future gains are in enterprise and code generation, which still have a massive runway for improvement. Consumer AI needs better integration, not just stronger models.

Bret Taylor explains the perception that AI progress has stalled. While improvements for casual tasks like trip planning are marginal, the reasoning capabilities of newer models have dramatically improved for complex work like software development or proving mathematical theorems.

AI Model Gains Are Now Incremental, Like a PhD Getting an 'A' Versus a 'B' | RiffOn