Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

The key uncertainty splitting AI bulls and bears isn't about data or compute, but 'spillover.' Will training models to be superhuman coders (a clean, verifiable task) also make them proficient at messy, open-ended tasks like business strategy? The size of this spillover effect is unknown and is the biggest determinant of AGI timelines.

Related Insights

Even a specialized task like coding involves a wide range of human-like interaction: brainstorming, searching, and more. This "AGI-completeness" means a powerful general model with a good "bedside manner" can outperform a narrowly specialized one, complicating the strategy for vertical AI apps.

Specialized coding models often fail because a developer's workflow isn't just writing code; it's a complex conversation involving brainstorming, compliance, and web research. The best coding assistants are the most generalist models because every complex task has AGI-like qualities.

Beyond enterprise sales, the intense focus on creating AI that can code is driven by a strategic belief that this is the most direct path to Artificial General Intelligence (AGI). Leaders like Anthropic believe an AI that can recursively improve its own code will be the first to achieve superintelligence.

Instead of a single, generalizable AI, we are creating 'Functional AGI'—a collection of specialized AIs layered together. This system will feel like AGI to users but lacks true cross-domain reasoning, as progress in one area (like coding) doesn't translate to others (like history).

The massive investment in AI coding tools isn't just about developer productivity. It's a strategic race based on the belief that an AI that can perfectly write and improve code is the key to achieving recursive self-improvement and, ultimately, AGI.

Current AI models resemble a student who grinds 10,000 hours on a narrow task. They achieve superhuman performance on benchmarks but lack the broad, adaptable intelligence of someone with less specific training but better general reasoning. This explains the gap between eval scores and real-world utility.

The ability to code is not just another domain for AI; it's a meta-skill. An AI that can program can build tools on demand to solve problems in nearly any digital domain, effectively simulating general competence. This makes mastery of code a form of instrumental, functional AGI for most economically valuable work.

The current focus on pre-training AI with specific tool fluencies overlooks the crucial need for on-the-job, context-specific learning. Humans excel because they don't need pre-rehearsal for every task. This gap indicates AGI is further away than some believe, as true intelligence requires self-directed, continuous learning in novel environments.

The path to AGI won't be uniform. Instead, we'll see 'jagged superintelligence,' where models achieve superhuman capabilities in specific verticals with high verifiability, such as coding, finance, and scientific research. These specialized peaks of excellence will appear long before a generalized intelligence is achieved.

Replit's CEO argues that today's LLMs are asymptoting on general reasoning tasks. Progress continues only in domains with binary outcomes, like coding, where synthetic data can be generated infinitely. This indicates a fundamental limitation of the current 'ingest the internet' approach for achieving AGI.