Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

The performance race in frontier AI models is irrelevant for most business use cases. The vast majority of enterprise AI traffic—an estimated 90%—will run on cheaper, older, or specialized open-source models that are sufficient for day-to-day operational tasks, rather than costly state-of-the-art ones.

Related Insights

Faced with rising costs from proprietary labs, sophisticated enterprise clients are building internal evaluation and routing systems. This allows them to use cheaper, open-source models for less complex tasks, optimizing for both cost and performance.

Releases like Cognition's SWE 2 and DeepSeek's V4.1 Flash show a mature market trend: optimizing for cost and efficiency over chasing absolute best performance. These models offer near-frontier capability on specific tasks at a fraction of the cost, enabling businesses to build sustainable, scalable AI features without exorbitant expenses.

With frontier models costing over 100x more than competent alternatives ($56 vs. 50¢ per million tokens), companies are burning cash. An estimated 98% of tasks sent to top-tier models don't require that power, an inefficiency driven by engineers who are disconnected from cost implications.

Large enterprises like AT&T manage soaring AI costs with a tiered strategy. They aim to use cheaper open-source models for 60-70% of internal tasks, keeping spending on expensive frontier models flat while overall AI usage grows. This treats premium models as specialized tools, not defaults.

Glean's co-founder argues that most enterprise tasks don't require expensive frontier models. Open-source alternatives are now capable enough for the vast majority of use cases. The primary adoption driver has shifted from data privacy to pure cost savings, as enterprises seek to control skyrocketing AI bills.

The era of using the most powerful AI model for every task is ending. Companies are now focused on the trade-off between quality, cost, and latency. The key question is no longer "Which model is best?" but "Which model is good enough for this task at the lowest price point?"

For typical enterprise tasks like code migration, using an optimized control plane with an open-source model can be over 16 times cheaper than using a frontier model like Claude Opus. While it may be slower, the massive cost savings make it a compelling business alternative.

An RBC analyst predicts an "80/20 world" for AI, where 80% of workloads can be handled by older, cheaper, or open-source models. However, the largest portion of the total addressable market (TAM) in terms of dollars will remain concentrated in the 20% of complex tasks that require cutting-edge frontier models.

Companies like Meta and Ramp are building AI routers to automatically send simple tasks to cheaper models. This trend shows the enterprise AI market is maturing past a 'one-model-fits-all' approach, focusing instead on cost management and operational efficiency by treating models as a commodity portfolio.

As enterprises scale AI, the high inference costs of frontier models become prohibitive. The strategic trend is to use large models for novel tasks, then shift 90% of recurring, common workloads to specialized, cost-effective Small Language Models (SLMs). This architectural shift dramatically improves both speed and cost.

90% of Enterprise AI Workloads Will Run on Cost-Efficient 'Good Enough' Models | RiffOn