We scan new podcasts and send you the top 5 insights daily.
Contrary to relying on a single frontier model, companies in production use a diverse portfolio of, on average, 32 different models. They switch between them to optimize for cost and performance on specific tasks, fueled by the rise of capable open-weight models.
The future of enterprise AI isn't choosing one provider. Instead, companies will use a "composable model" approach, routing queries to a combination of powerful frontier models and their own fine-tuned open-source models. This strategy, dubbed the "council of LLMs," optimizes for cost, performance, and specialization on proprietary data.
The era of relying on a single frontier AI model is ending. A combination of factors—the high cost of agentic workloads, compute shortages, and government intervention seen with Fable 5—is pushing businesses toward multi-model architectures to optimize for cost, speed, and resilience.
Just as developers use various databases for different needs, AI applications will rely on a "constellation" of specialized models. Some tasks will require expensive, high-reasoning models, while others will prioritize low-latency or low-cost models. The market will become heterogeneous, not monolithic.
Instead of relying on one powerful model for all tasks, the leading strategy is 'smart routing'—using a panel of models and directing each task to the most appropriate one. This compound architecture demonstrably beats single frontier models on both cost and performance.
The AI model landscape isn't a simple ladder of best to worst. Instead, it's a "spiky" frontier where different models offer unique strengths. For example, one model may excel at complex, niche problems while another is faster, more affordable, and better for collaborative, general-purpose tasks, necessitating a multi-tool approach.
An intelligent AI orchestration layer can achieve a cost-to-accuracy balance superior to any single model. By routing queries to a portfolio of different models (large, small, specialized), it creates a new Pareto frontier, delivering higher success rates at a lower average cost than relying on one "best" model.
Initially, even OpenAI believed a single, ultimate 'model to rule them all' would emerge. This thinking has completely changed to favor a proliferation of specialized models, creating a healthier, less winner-take-all ecosystem where different models serve different needs.
Instead of relying on a single large AI model, companies are adopting "model orchestration" to control costs. This involves using a router to send prompts to the most appropriate model based on the task, often cascading from cheap, small models to more expensive ones only when necessary.
The belief that a single, god-level foundation model would dominate has proven false. Horowitz points to successful AI applications like Cursor, which uses 13 different models. This shows that value lies in the complex orchestration and design at the application layer, not just in having the largest single model.
The most advanced AI users are 'polyamorous' with models, using an average of 3.5 different tools. This indicates a mature usage pattern where users select the best model for a specific job rather than relying on a single, all-purpose AI, challenging the 'winner-take-all' market theory.