We scan new podcasts and send you the top 5 insights daily.
Agentic AI has a different computational profile than previous generative AI. It uses much larger inputs with high reusability, creating larger KV caches. This change makes GPU architectures, which were optimized for earlier workloads, inefficient for the new demands of agentic inference.
Early AI models were compute-heavy with little memory. The next evolution, driven by agentic AI, requires massive memory stores, mirroring the human brain's structure. This shift is fueling the "Rampocalypse" and will make memory as critical as compute.
The rise of agentic AI, which runs multiple parallel processes, is elevating the CPU's role from a secondary component to a critical 'conductor' for orchestrating GPU tasks, creating new demand and design considerations.
While GPUs are key for model training, the next AI wave of autonomous agents relies more on CPUs. The task of controlling and orchestrating multiple agents and tool calls is fundamentally a CPU-based process. This is creating a new hardware bottleneck and shifting focus to CPU manufacturers.
The current AI boom focuses on GPUs for "thinking" (Gen AI). The next phase, "Agentic AI" for "doing," will rely heavily on CPUs for task orchestration and memory for context, creating new investment opportunities in this previously overshadowed hardware.
The era of dual-purpose AI chips is ending. The overwhelming demand for real-time processing from AI agents is forcing companies like Google and NVIDIA to create dedicated, inference-optimized hardware. This marks a fundamental and permanent split in the AI infrastructure market, separating training from inference.
Contrary to the idea that infrastructure problems get commoditized, AI inference is growing more complex. This is driven by three factors: (1) increasing model scale (multi-trillion parameters), (2) greater diversity in model architectures and hardware, and (3) the shift to agentic systems that require managing long-lived, unpredictable state.
Mark, CTO of AMD, states that the explosion of agentic AI workflows has created an unforeseen demand for a balanced compute architecture. These complex, multi-step processes require a CPU to GPU ratio approaching 1:1, a significant shift from traditional GPU-heavy AI training and inference models.
The shift from simple query-based AI to agentic AI, where AI calls itself recursively to solve complex tasks, increases compute demand by orders of magnitude. Most people, especially non-coders, fail to grasp this exponential shift, leading them to consistently underestimate the scale and duration of the AI infrastructure build-out.
Jensen Huang quantifies the massive computational leap required for advanced AI. The move from generative AI to reasoning was a 100x compute increase, and the subsequent move to agentic systems that can perform work represents another 100x jump. This results in a staggering 10,000x increase in computational demand in just two years.
GPUs are ill-suited for generating output tokens because they process models in discrete chunks called "kernels." This method requires constant data transfer to and from external memory, creating a significant bottleneck limited by memory bandwidth, which ultimately slows down real-time inference.