We scan new podcasts and send you the top 5 insights daily.
Clay CEO Kareem Amin is less focused on the ratio of open vs. closed-source model usage. His team's priority is a pragmatic engineering approach: if a task is repeatable, they write code to automate it and avoid model inference costs entirely. For other tasks, an internal router selects the most efficient model based on cost, speed, or quality.
Faced with rising costs from proprietary labs, sophisticated enterprise clients are building internal evaluation and routing systems. This allows them to use cheaper, open-source models for less complex tasks, optimizing for both cost and performance.
Companies like legal AI provider Lagora don't rely on a single frontier model. Instead, they build their own internal routers that intelligently direct different tasks to the most suitable model—whether it's from OpenAI, Anthropic, or open-source. This allows them to optimize for performance, cost, and specific capabilities for each component of their workflow.
They built an internal system that routes AI tasks to the most appropriate model, favoring cheaper, self-hosted open-weight models for 99% of requests. This dramatically cuts costs and prevents dependency on any single frontier model provider.
Sophisticated startups are adopting a hybrid AI strategy, using expensive frontier models for complex work while routing routine tasks like data extraction to cheaper open-source alternatives. This workload routing enables them to reduce costs by 5 to 20 times, creating more sustainable business models.
The optimal strategy for enterprise AI is not to rely solely on expensive frontier models. Instead, companies use a powerful model like Claude or GPT-4 to plan tasks and then delegate the execution to cheaper, fine-tuned open-source models. This massively reduces cost while maintaining high performance.
Glean's co-founder argues that most enterprise tasks don't require expensive frontier models. Open-source alternatives are now capable enough for the vast majority of use cases. The primary adoption driver has shifted from data privacy to pure cost savings, as enterprises seek to control skyrocketing AI bills.
For typical enterprise tasks like code migration, using an optimized control plane with an open-source model can be over 16 times cheaper than using a frontier model like Claude Opus. While it may be slower, the massive cost savings make it a compelling business alternative.
A cost-saving workflow is emerging where developers use expensive frontier models for high-level "thinking" and planning stages of a complex task. Once the plan is established, the more routine and high-volume execution steps are routed to cheaper, often open-source, models to optimize both performance and cost.
A sophisticated gateway that routes queries to different models based on complexity is key to managing AI costs. Simple tasks go to cheap, open-source models, while difficult ones use the frontier. This "expert pattern" allows token usage to rise while keeping costs flat.
Companies like Meta and Ramp are building AI routers to automatically send simple tasks to cheaper models. This trend shows the enterprise AI market is maturing past a 'one-model-fits-all' approach, focusing instead on cost management and operational efficiency by treating models as a commodity portfolio.