We scan new podcasts and send you the top 5 insights daily.
The amount of compute spent on web search for an AI agent should be proportional to the cost of the LLM it's feeding. Expensive models warrant more pre-processing on the search side to optimize their costly context windows, while cheaper models do not.
Enterprises are currently overspending on tokens by sending all queries to the most powerful LLMs. A new software category will emerge to intelligently route requests to smaller, cheaper models when possible, creating a critical efficiency and cost-saving layer between companies and foundational model providers.
AI's hunger for context is making search a critical but expensive component. As illustrated by Turbo Puffer's origin, a single recommendation feature using vector embeddings can cost tens of thousands per month, forcing companies to find cheaper solutions to make AI features economically viable at scale.
The growth of LLM context windows has stalled not primarily due to technical barriers, but because multi-million token requests can cost users several dollars per query, leading to low demand. The industry is shifting focus to "smart context" techniques like compaction and retrieval to provide relevant information without the prohibitive cost of massive context.
When an AI agent performs web searches, it generates multiple queries for a single task. These different queries often lead to the same URLs, causing the agent to revisit and process the same content repeatedly, dramatically increasing token consumption and cost.
Companies are building intelligent systems that analyze a user's prompt and automatically route it to the most cost-effective model that can handle the task. This avoids using expensive frontier models for simple requests, with some companies like Coinbase successfully keeping costs flat despite exponential usage growth.
A sophisticated gateway that routes queries to different models based on complexity is key to managing AI costs. Simple tasks go to cheap, open-source models, while difficult ones use the frontier. This "expert pattern" allows token usage to rise while keeping costs flat.
Advanced agentic memory can act as a cache for LLM-generated answers. For similar queries, an agent can retrieve a cached response via vector search and validate it with a cheap evaluative LLM. This avoids expensive generative calls, combating “token maxing” and preventing inconsistent answers.
A production AI agent performs tasks of varying difficulty. Forcing all requests through a single, expensive frontier model is inefficient. A better architecture routes tasks to the most appropriate model: small, cheap open models for high-volume, low-difficulty work like retrieval, reserving the costly frontier API only for high-stakes reasoning where it matters.
For difficult, multi-step tasks, a more capable LLM can reach a solution with fewer iterations than a smaller model. Despite a higher per-token price, this efficiency can lead to a lower total token count and a cheaper overall cost for the task, proving that cheaper-per-token isn't always cheaper-per-task.
Current web search API prices are too high for the coming wave of AI agents using cheap models. Spending 80-90% of a task's cost on search for a cheap model is 'entirely silly.' The market must race to the bottom on price to enable the 1000x scale increase.