Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

As AI models evolve to mirror the human brain, their memory requirements are skyrocketing, creating a 'RAMpocalypse.' The industry's focus will shift from being purely compute-centric to a dual focus on memory and compute, making high-bandwidth memory a critical and scarce resource.

Related Insights

The demand for HBM memory for AI is causing a global shortage because of a ~4:1 manufacturing trade-off: each bit of HBM produced consumes capacity that could have made four bits of standard DRAM. This supply crunch will raise prices for all electronics, from phones to PCs.

AI workloads are limited by memory bandwidth, not capacity. While commodity DRAM offers more bits per wafer, its bandwidth is over an order of magnitude lower than specialized HBM. This speed difference would starve the GPU's compute cores, making the extra capacity useless and creating a massive performance bottleneck.

Unlike past tech cycles with a single constraint, the AI boom is constrained by numerous interdependent bottlenecks at once: power, transmission, memory, optical components, and skilled labor. Solving one piece (e.g., memory supply) doesn't fix the overall systems-level challenge, making the problem uniquely complex.

The AI industry's growth constraint is a swinging pendulum. While power and data center space are the current bottlenecks (2024-25), the energy supply chain is diverse. By 2027, the bottleneck will revert to semiconductor manufacturing, as leading-edge fab capacity (e.g., TSMC, HBM memory) is highly concentrated and takes years to expand.

Hardware shortages act as a catalyst for software innovation. The 'Kimi moment,' where a Chinese model introduced major memory efficiency improvements, demonstrates a recurring pattern: when a component like memory becomes a bottleneck, the ecosystem responds with algorithmic breakthroughs to reduce demand for it.

The focus in AI has evolved from rapid software capability gains to the physical constraints of its adoption. The demand for compute power is expected to significantly outstrip supply, making infrastructure—not algorithms—the defining bottleneck for future growth.

While NVIDIA's GPUs have been the primary AI constraint, the bottleneck is now moving to other essential subsystems. Memory, networking interconnects, and power management are emerging as the next critical choke points, signaling a new wave of investment opportunities in the hardware stack beyond core compute.

The current AI boom focuses on GPUs for "thinking" (Gen AI). The next phase, "Agentic AI" for "doing," will rely heavily on CPUs for task orchestration and memory for context, creating new investment opportunities in this previously overshadowed hardware.

The primary bottleneck for AI inference is now memory (HBM), not compute. To circumvent this, industry giants Nvidia and AWS are making multi-billion dollar deals for systems from Groq and Cerebrus that use on-chip SRAM, which is faster and not subject to the same supply constraints.

While GPUs are the current focus, the rising cost of memory (DRAM) is creating a massive incentive for a disruptive innovation. This makes the memory complex, not models, a likely area for China's next big AI breakthrough, as it seeks to widen the bottleneck.

The Next AI Bottleneck is Memory, Leading to a 'RAMpocalypse' | RiffOn