Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Companies no longer chase the single most powerful AI model. The new standard is creating a sophisticated architecture of multiple models, matching the right tool to the right task based on capability, efficiency, and cost, which allows for greater optimization across the enterprise.

Related Insights

The future of enterprise AI isn't choosing one provider. Instead, companies will use a "composable model" approach, routing queries to a combination of powerful frontier models and their own fine-tuned open-source models. This strategy, dubbed the "council of LLMs," optimizes for cost, performance, and specialization on proprietary data.

The era of using the most powerful AI model for every task is ending. Companies are now focused on the trade-off between quality, cost, and latency. The key question is no longer "Which model is best?" but "Which model is good enough for this task at the lowest price point?"

As AI use matures, the critical task is no longer just picking the best model. It's building a sophisticated internal architecture—including routers, monitors, and guardrails—to manage costs and route tasks effectively, treating AI as a system to be engineered.

According to Meta's CTO, the era of one monolithic model doing everything is over. The current frontier involves using a 'harness' that intelligently routes tasks to a collection of different, specialized models based on cost, latency, and capability.

Instead of relying on one powerful model for all tasks, the leading strategy is 'smart routing'—using a panel of models and directing each task to the most appropriate one. This compound architecture demonstrably beats single frontier models on both cost and performance.

Enterprises will shift from relying on a single large language model to using orchestration platforms. These platforms will allow them to 'hot swap' various models—including smaller, specialized ones—for different tasks within a single system, optimizing for performance, cost, and use case without being locked into one provider.

The future of enterprise AI isn't a winner-take-all model. Instead, companies will use a mix: cheap open-weight models for routine tasks and premium, specialized models for critical functions like genomics. Cloud providers offering this "mixture of models" will have a strategic advantage over pure-play model providers.

Contrary to relying on a single frontier model, companies in production use a diverse portfolio of, on average, 32 different models. They switch between them to optimize for cost and performance on specific tasks, fueled by the rise of capable open-weight models.

Companies like Meta and Ramp are building AI routers to automatically send simple tasks to cheaper models. This trend shows the enterprise AI market is maturing past a 'one-model-fits-all' approach, focusing instead on cost management and operational efficiency by treating models as a commodity portfolio.

According to NVIDIA's VP, the modern approach to enterprise AI involves mixing models. Use expensive, powerful frontier models for complex, high-value tasks like agentic planning. For more trivial, high-volume tasks like document summarization, use cheaper, fine-tuned open-source models to optimize cost.