Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

The future of enterprise AI isn't a winner-take-all model. Instead, companies will use a mix: cheap open-weight models for routine tasks and premium, specialized models for critical functions like genomics. Cloud providers offering this "mixture of models" will have a strategic advantage over pure-play model providers.

Related Insights

The future of enterprise AI isn't choosing one provider. Instead, companies will use a "composable model" approach, routing queries to a combination of powerful frontier models and their own fine-tuned open-source models. This strategy, dubbed the "council of LLMs," optimizes for cost, performance, and specialization on proprietary data.

Early enterprise AI adoption mirrored the initial, inefficient use of AWS, with rampant experimentation. Now, companies are maturing, learning to apply AI strategically, much like a savvy Costco shopper who targets specific items instead of wandering every aisle. This shift involves using cheaper or open-source models for simpler tasks and reserving frontier models for high-value problems.

Just as developers use various databases for different needs, AI applications will rely on a "constellation" of specialized models. Some tasks will require expensive, high-reasoning models, while others will prioritize low-latency or low-cost models. The market will become heterogeneous, not monolithic.

Enterprises will shift from relying on a single large language model to using orchestration platforms. These platforms will allow them to 'hot swap' various models—including smaller, specialized ones—for different tasks within a single system, optimizing for performance, cost, and use case without being locked into one provider.

An intelligent AI orchestration layer can achieve a cost-to-accuracy balance superior to any single model. By routing queries to a portfolio of different models (large, small, specialized), it creates a new Pareto frontier, delivering higher success rates at a lower average cost than relying on one "best" model.

Contrary to relying on a single frontier model, companies in production use a diverse portfolio of, on average, 32 different models. They switch between them to optimize for cost and performance on specific tasks, fueled by the rise of capable open-weight models.

As AI token consumption becomes a major budget item, companies are moving beyond using a single frontier model. Every organization will need a portfolio of models, including cheaper options for less complex tasks, to manage the "madness" of runaway costs.

Companies like Meta and Ramp are building AI routers to automatically send simple tasks to cheaper models. This trend shows the enterprise AI market is maturing past a 'one-model-fits-all' approach, focusing instead on cost management and operational efficiency by treating models as a commodity portfolio.

The greatest value in AI won't be captured by frontier labs alone. Instead, companies in the "applied layer" are incentivized to build routing systems that use expensive frontier models for high-level orchestration while deploying cheaper open-source models for bulk tasks, creating a more efficient, barbell-shaped cost structure.

According to NVIDIA's VP, the modern approach to enterprise AI involves mixing models. Use expensive, powerful frontier models for complex, high-value tasks like agentic planning. For more trivial, high-volume tasks like document summarization, use cheaper, fine-tuned open-source models to optimize cost.