We scan new podcasts and send you the top 5 insights daily.
According to NVIDIA's VP, the modern approach to enterprise AI involves mixing models. Use expensive, powerful frontier models for complex, high-value tasks like agentic planning. For more trivial, high-volume tasks like document summarization, use cheaper, fine-tuned open-source models to optimize cost.
The future of enterprise AI isn't choosing one provider. Instead, companies will use a "composable model" approach, routing queries to a combination of powerful frontier models and their own fine-tuned open-source models. This strategy, dubbed the "council of LLMs," optimizes for cost, performance, and specialization on proprietary data.
Don't use the most powerful and expensive AI model for every task. Use cheaper, faster models like Anthropic's Haiku for high-volume, simple jobs and reserve powerful models like Opus for complex reasoning. This strategy can reduce costs by over 99%, turning a potential $150 task into a $1.50 one.
To combat rising AI costs, firms are creating hybrid systems that use cheaper "worker" models for routine tasks while delegating complex problems to powerful "advisor" models. This approach, used by Harvey and explored by Microsoft, can outperform state-of-the-art models alone for a fraction of the cost.
Instead of relying on one powerful model for all tasks, the leading strategy is 'smart routing'—using a panel of models and directing each task to the most appropriate one. This compound architecture demonstrably beats single frontier models on both cost and performance.
An intelligent AI orchestration layer can achieve a cost-to-accuracy balance superior to any single model. By routing queries to a portfolio of different models (large, small, specialized), it creates a new Pareto frontier, delivering higher success rates at a lower average cost than relying on one "best" model.
Contrary to relying on a single frontier model, companies in production use a diverse portfolio of, on average, 32 different models. They switch between them to optimize for cost and performance on specific tasks, fueled by the rise of capable open-weight models.
As AI token consumption becomes a major budget item, companies are moving beyond using a single frontier model. Every organization will need a portfolio of models, including cheaper options for less complex tasks, to manage the "madness" of runaway costs.
Companies like Meta and Ramp are building AI routers to automatically send simple tasks to cheaper models. This trend shows the enterprise AI market is maturing past a 'one-model-fits-all' approach, focusing instead on cost management and operational efficiency by treating models as a commodity portfolio.
As enterprises scale AI, the high inference costs of frontier models become prohibitive. The strategic trend is to use large models for novel tasks, then shift 90% of recurring, common workloads to specialized, cost-effective Small Language Models (SLMs). This architectural shift dramatically improves both speed and cost.
The smartest 'AI-pilled' companies adopt a two-tiered model strategy. They use expensive, frontier models for internal, high-leverage tasks like creating new knowledge and optimizing processes. However, they use cheaper, open-weight models in the 'bill of materials' for the customer-facing product to manage costs effectively.