Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

A sophisticated gateway that routes queries to different models based on complexity is key to managing AI costs. Simple tasks go to cheap, open-source models, while difficult ones use the frontier. This "expert pattern" allows token usage to rise while keeping costs flat.

Related Insights

Faced with rising costs from proprietary labs, sophisticated enterprise clients are building internal evaluation and routing systems. This allows them to use cheaper, open-source models for less complex tasks, optimizing for both cost and performance.

Sophisticated startups are adopting a hybrid AI strategy, using expensive frontier models for complex work while routing routine tasks like data extraction to cheaper open-source alternatives. This workload routing enables them to reduce costs by 5 to 20 times, creating more sustainable business models.

Don't use the most powerful and expensive AI model for every task. Use cheaper, faster models like Anthropic's Haiku for high-volume, simple jobs and reserve powerful models like Opus for complex reasoning. This strategy can reduce costs by over 99%, turning a potential $150 task into a $1.50 one.

To manage AI costs effectively, companies should avoid simply capping token usage, as this kills innovation. A better strategy is to build intelligent routers that assess a task's complexity and dynamically route it to the most appropriate model—powerful models for hard tasks, cheaper ones for simple tasks.

To optimize AI costs and sustainability, UBS employs a "model garden" with various frontier and smaller models. An internal AI system then routes employee questions to the most appropriate, cost-effective model, preventing the wasteful use of powerful, expensive LLMs for simple, non-frontier problems.

Companies are building intelligent systems that analyze a user's prompt and automatically route it to the most cost-effective model that can handle the task. This avoids using expensive frontier models for simple requests, with some companies like Coinbase successfully keeping costs flat despite exponential usage growth.

To prevent AI agent usage costs from spiraling, GitHub expects the solution will be intelligent model routing. These systems will automatically select the most efficient and cost-effective AI model for a given task, such as using a cheap model for simple refactoring instead of a powerful, expensive one.

A production AI agent performs tasks of varying difficulty. Forcing all requests through a single, expensive frontier model is inefficient. A better architecture routes tasks to the most appropriate model: small, cheap open models for high-volume, low-difficulty work like retrieval, reserving the costly frontier API only for high-stakes reasoning where it matters.

To control inference costs, companies are implementing model routing systems. They differentiate between expensive tokens from frontier models for complex reasoning and cheaper tokens from fine-tuned open-source models for simpler workflow tasks. This tiered approach optimizes both performance and budget, avoiding "token maxing."

An optimal AI architecture routes tasks to different models based on complexity and risk. Simple, low-stakes work like data extraction should go to the cheapest models. Ambiguous, high-stakes work like system design warrants expensive frontier models, where preventing one engineering mistake justifies the premium token cost.