Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Frontier model providers like OpenAI are not incentivized to reduce your token costs, creating lock-in. Independent agent labs (e.g., Cognition) act like brokers, routing tasks to the most efficient models, which is a more sustainable and affordable long-term strategy for building your company.

Related Insights

Enterprises are currently overspending on tokens by sending all queries to the most powerful LLMs. A new software category will emerge to intelligently route requests to smaller, cheaper models when possible, creating a critical efficiency and cost-saving layer between companies and foundational model providers.

To defend against large model providers, AI coding startups like Cognition are moving from being "model neutral" to "agent neutral." They now integrate competing coding agents (e.g., Claude Code) into their platforms, shifting their value proposition to being the essential workflow and orchestration layer for developers.

For most startups, training a custom foundation model is a waste of capital. The winning strategy is to focus on workflow and proprietary data, building a "headless" product that uses a model router to switch between the cheapest, most effective LLMs for any given task.

With 80% of revenue tied to token usage, leading model providers are not incentivized to offer features like auto-routing to cheaper models. This business model conflict creates a competitive vulnerability and an opportunity for third-party tools like Cursor to win by optimizing developer experience and cost-efficiency.

To manage AI costs effectively, companies should avoid simply capping token usage, as this kills innovation. A better strategy is to build intelligent routers that assess a task's complexity and dynamically route it to the most appropriate model—powerful models for hard tasks, cheaper ones for simple tasks.

Large enterprises are avoiding commitment to a single AI provider like OpenAI or Anthropic. Instead, they're building control planes and abstraction layers that allow them to hot-swap the underlying models, mitigating technology risk and preventing dependence on one provider's terms of service.

In response to budget blowouts from agentic AI, enterprises are moving beyond simple adoption to active cost management. A new "token efficiency" stack is emerging, featuring tactics like model routing to cheaper alternatives (e.g., DeepSeek) and custom post-trained models to reduce reliance on expensive foundation models.

To control inference costs, companies are implementing model routing systems. They differentiate between expensive tokens from frontier models for complex reasoning and cheaper tokens from fine-tuned open-source models for simpler workflow tasks. This tiered approach optimizes both performance and budget, avoiding "token maxing."

Relying on a single foundation model provider is inefficient, as different models excel at different tasks. An independent, third-party agent platform is crucial to act as a router, selecting the optimal model for each job, thereby maximizing performance while controlling spiraling inference costs for enterprises.

Enterprises shouldn't lock into a single AI lab like OpenAI or Anthropic. Instead, they need a multi-model strategy using evals and routing to leverage different models for different tasks, ensuring they aren't trapped when the 'seasons' inevitably change.

Build Your AI Software Factory on Independent Agent Labs, Not Frontier Model Platforms | RiffOn