Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

They built an internal system that routes AI tasks to the most appropriate model, favoring cheaper, self-hosted open-weight models for 99% of requests. This dramatically cuts costs and prevents dependency on any single frontier model provider.

Related Insights

Faced with rising costs from proprietary labs, sophisticated enterprise clients are building internal evaluation and routing systems. This allows them to use cheaper, open-source models for less complex tasks, optimizing for both cost and performance.

Enterprises are currently overspending on tokens by sending all queries to the most powerful LLMs. A new software category will emerge to intelligently route requests to smaller, cheaper models when possible, creating a critical efficiency and cost-saving layer between companies and foundational model providers.

Companies like legal AI provider Lagora don't rely on a single frontier model. Instead, they build their own internal routers that intelligently direct different tasks to the most suitable model—whether it's from OpenAI, Anthropic, or open-source. This allows them to optimize for performance, cost, and specific capabilities for each component of their workflow.

Sophisticated startups are adopting a hybrid AI strategy, using expensive frontier models for complex work while routing routine tasks like data extraction to cheaper open-source alternatives. This workload routing enables them to reduce costs by 5 to 20 times, creating more sustainable business models.

The optimal strategy for enterprise AI is not to rely solely on expensive frontier models. Instead, companies use a powerful model like Claude or GPT-4 to plan tasks and then delegate the execution to cheaper, fine-tuned open-source models. This massively reduces cost while maintaining high performance.

Enterprises will shift from relying on a single large language model to using orchestration platforms. These platforms will allow them to 'hot swap' various models—including smaller, specialized ones—for different tasks within a single system, optimizing for performance, cost, and use case without being locked into one provider.

Instead of relying on a single large AI model, companies are adopting "model orchestration" to control costs. This involves using a router to send prompts to the most appropriate model based on the task, often cascading from cheap, small models to more expensive ones only when necessary.

Bending Spoons avoids massive token costs by using self-hosted, narrow-purpose AI models for specific tasks. An internal AI orchestrator routes jobs to the most cost-effective model, reserving expensive frontier models only for complex tasks or supervision. This strategy allows them to generate over 90% of their code with AI at a modest cost.

Companies are building intelligent systems that analyze a user's prompt and automatically route it to the most cost-effective model that can handle the task. This avoids using expensive frontier models for simple requests, with some companies like Coinbase successfully keeping costs flat despite exponential usage growth.

A production AI agent performs tasks of varying difficulty. Forcing all requests through a single, expensive frontier model is inefficient. A better architecture routes tasks to the most appropriate model: small, cheap open models for high-volume, low-difficulty work like retrieval, reserving the costly frontier API only for high-stakes reasoning where it matters.

Bending Spoons' AI Orchestration Layer Avoids Vendor Lock-In by Using Cheaper, Self-Hosted Models | RiffOn