Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

SAP CEO Christian Klein reveals that his teams constantly test and switch the underlying AI models for their agents. This "model switching" has become a core strategy to manage runaway token costs, prioritizing the best price-to-outcome ratio over allegiance to a single frontier model.

Related Insights

The era of using the most powerful AI model for every task is ending. Companies are now focused on the trade-off between quality, cost, and latency. The key question is no longer "Which model is best?" but "Which model is good enough for this task at the lowest price point?"

The most sophisticated AI users aren't locking into one provider. Faced with a 13x annual increase in token costs, they leverage multiple models and routing platforms like OpenRouter to optimize for price and performance. This behavior suggests a future of model commoditization, not monopoly.

For enterprises using AI at scale, the most impactful cost-saving measure is not just smart routing but aggressively shifting workloads to newer, more efficient models as they are released. This constant deflationary pressure from model innovation provides significant savings without requiring changes to user behavior.

In response to budget blowouts from agentic AI, enterprises are moving beyond simple adoption to active cost management. A new "token efficiency" stack is emerging, featuring tactics like model routing to cheaper alternatives (e.g., DeepSeek) and custom post-trained models to reduce reliance on expensive foundation models.

As AI token consumption becomes a major budget item, companies are moving beyond using a single frontier model. Every organization will need a portfolio of models, including cheaper options for less complex tasks, to manage the "madness" of runaway costs.

Large customers are aggressively optimizing AI spend by abandoning a one-size-fits-all frontier model approach. One software provider is saving nearly $700,000 annually by switching to a much cheaper OpenAI model for a high-volume task, signaling a market-wide shift towards cost-efficiency and model routing.

Companies no longer chase the single most powerful AI model. The new standard is creating a sophisticated architecture of multiple models, matching the right tool to the right task based on capability, efficiency, and cost, which allows for greater optimization across the enterprise.

As AI costs rise, using one powerful frontier model for every task is no longer financially viable. The solution is to create a dedicated "Model Sommelier" role responsible for curating a portfolio of models, continuously testing and selecting the most cost-effective option for each specific business use case.

A few months ago, the fear was AI replacing SaaS businesses. Now, the pressing issue is managing massive AI bills. This has elevated 'token economics'—optimizing costs by using different models for different tasks (model routing)—from an advanced technique to a non-negotiable, table-stakes practice for any serious AI implementation.

Early enterprise AI adoption featured 'token maxing'—unrestricted use of expensive models. The trend is now 'token efficiency' via smart routing platforms that delegate low-value tasks to cheaper models. This substitution optimizes costs and puts margin pressure on premium frontier models.

Enterprises Now Constantly Switch AI Models to Optimize Cost-to-Performance | RiffOn