We scan new podcasts and send you the top 5 insights daily.
The speaker found most users abandoned slow (6-8 second) searches before results loaded. Implementing a cache to reduce load times to 0.2 seconds had a greater impact on conversions than any improvements to the LLM's result quality. Speed was more critical than intelligence.
Analysis of AI spending shows users will pay significantly more for faster model inference (e.g., 6x price for 2x speed), prioritizing interactivity over marginal gains in intelligence. This mirrors how e-commerce conversions are highly sensitive to latency, suggesting speed is a critical, high-value feature for AI products.
Criteo has just milliseconds to respond to an ad request. This extreme speed requirement dictates their AI architecture, forcing them to pre-compute and cache user and product embeddings. Real-time inference is limited to fast operations with only marginal updates for the user's latest action.
As frontier AI models reach a plateau of perceived intelligence, the key differentiator is shifting to user experience. Low-latency, reliable performance is becoming more critical than marginal gains on benchmarks, making speed the next major competitive vector for AI products like ChatGPT.
In 2001, Google realized its combined server RAM could hold a full copy of its web index. Moving from disk-based to in-memory systems eliminated slow disk seeks, enabling complex queries with synonyms and semantic expansion. This fundamentally improved search quality long before LLMs became mainstream.
Google's focus on fast, cost-effective models like Gemini 3.5 Flash is driven by the needs of its massive-scale products (e.g., Search). For billions of users, low latency and cost are more critical than absolute peak performance, as users are often unwilling to wait for a slightly smarter but slower response.
Frame the value of speed beyond just a better user experience. Ask customers how they could use the time saved by faster AI responses to pack in more value, create premium product tiers, or open entirely new revenue streams that were previously impossible.
A key way to improve consumer LLM speed and cost is to cache the results for frequently asked, static questions like "When was OpenAI founded?" This approach, similar to Google's knowledge panels, would provide instant answers for a large cohort of queries without engaging expensive GPU resources for every request.
Breaking from transformer dominance, Shopify leverages Liquid AI's state-space-like models for high-value tasks. For search query understanding, they run a 300M parameter Liquid model with an impressive 30ms end-to-end latency, a feat difficult to achieve with traditional architectures.
Pichai reveals Google's operational tactic for maintaining speed: teams have "latency budgets" in milliseconds. If a feature saves time, they earn a credit they can "spend" on new capabilities, ensuring the user experience remains fast while the product evolves.
Unlike streaming text from LLMs, image generation forces users to wait. An A/B test by one of Fal's customers proved that increased latency directly harms user engagement and the number of images created, much like slow page loads hurt e-commerce sales.