Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Fintech company Ramp is expanding into AI infrastructure by launching a 'model router.' This tool addresses growing CFO frustration with uncontrolled AI spending by intelligently routing tasks to the most cost-effective model. This move indicates that AI cost management is becoming a critical new product category for enterprise software.

Related Insights

Enterprises are currently overspending on tokens by sending all queries to the most powerful LLMs. A new software category will emerge to intelligently route requests to smaller, cheaper models when possible, creating a critical efficiency and cost-saving layer between companies and foundational model providers.

Contrary to the belief that enterprises have unlimited budgets, they are focused on the ROI of their AI spend. As agentic workflows cause token bills to skyrocket, orchestration tools that intelligently route queries to the most cost-effective model for a given task are becoming essential infrastructure.

To manage AI costs effectively, companies should avoid simply capping token usage, as this kills innovation. A better strategy is to build intelligent routers that assess a task's complexity and dynamically route it to the most appropriate model—powerful models for hard tasks, cheaper ones for simple tasks.

Instead of relying on a single large AI model, companies are adopting "model orchestration" to control costs. This involves using a router to send prompts to the most appropriate model based on the task, often cascading from cheap, small models to more expensive ones only when necessary.

In response to budget blowouts from agentic AI, enterprises are moving beyond simple adoption to active cost management. A new "token efficiency" stack is emerging, featuring tactics like model routing to cheaper alternatives (e.g., DeepSeek) and custom post-trained models to reduce reliance on expensive foundation models.

Companies are building intelligent systems that analyze a user's prompt and automatically route it to the most cost-effective model that can handle the task. This avoids using expensive frontier models for simple requests, with some companies like Coinbase successfully keeping costs flat despite exponential usage growth.

The recent focus on model routers signals a maturation of enterprise AI strategy. The initial "growth at all costs" phase, which encouraged rampant employee use ("token maxing"), is giving way to a new era of cost optimization and demonstrating clear ROI on AI investments.

Large customers are aggressively optimizing AI spend by abandoning a one-size-fits-all frontier model approach. One software provider is saving nearly $700,000 annually by switching to a much cheaper OpenAI model for a high-volume task, signaling a market-wide shift towards cost-efficiency and model routing.

To prevent AI agent usage costs from spiraling, GitHub expects the solution will be intelligent model routing. These systems will automatically select the most efficient and cost-effective AI model for a given task, such as using a cheap model for simple refactoring instead of a powerful, expensive one.

As AI costs rise, using one powerful frontier model for every task is no longer financially viable. The solution is to create a dedicated "Model Sommelier" role responsible for curating a portfolio of models, continuously testing and selecting the most cost-effective option for each specific business use case.