We scan new podcasts and send you the top 5 insights daily.
For 20 years, chip design focused on scaling flops (a million-fold increase), while memory bandwidth saw only a 40x increase, creating a massive bottleneck for large AI models. Fractile's core bet is that prioritizing extremely high-bandwidth memory access is the key to unlocking the next level of AI performance and efficiency.
The "memory wall" is a growing chasm between compute power and memory access. In the last decade, GPU flops improved 120-fold, but memory bandwidth only increased 17-fold. This divergence makes memory-bound workloads like AI inference an increasingly severe bottleneck for modern hardware.
AI workloads are limited by memory bandwidth, not capacity. While commodity DRAM offers more bits per wafer, its bandwidth is over an order of magnitude lower than specialized HBM. This speed difference would starve the GPU's compute cores, making the extra capacity useless and creating a massive performance bottleneck.
The next wave of AI silicon may pivot from today's compute-heavy architectures to memory-centric ones optimized for inference. This fundamental shift would allow high-performance chips to be produced on older, more accessible 7-14nm manufacturing nodes, disrupting the current dependency on cutting-edge fabs.
As AI models evolve to mirror the human brain, their memory requirements are skyrocketing, creating a 'RAMpocalypse.' The industry's focus will shift from being purely compute-centric to a dual focus on memory and compute, making high-bandwidth memory a critical and scarce resource.
While NVIDIA's GPUs have been the primary AI constraint, the bottleneck is now moving to other essential subsystems. Memory, networking interconnects, and power management are emerging as the next critical choke points, signaling a new wave of investment opportunities in the hardware stack beyond core compute.
While many focus on compute metrics like FLOPS, the primary bottleneck for large AI models is memory bandwidth—the speed of loading weights into the GPU. This single metric is a better indicator of real-world performance from one GPU generation to the next than raw compute power.
There exists a "scaling law for bandwidth." By solving the memory bottleneck, chip designers enable new AI model architectures that are computationally cheaper. For example, making Mixture-of-Experts (MoE) models significantly sparser saves flops but is prohibitive on today's bandwidth-starved GPUs. High-bandwidth chips make these more efficient architectures viable.
While training AI models is a compute-bound problem where more flops yield better results, inference (running the model) is memory-bound. Each token generation requires reading all model weights from memory, making memory bandwidth, not raw processing power, the primary performance bottleneck.
The dominant AI strategy of building increasingly larger models is becoming unsustainable. The primary constraint is memory, which is described as "already broken." Consequently, leading companies are abandoning the "scaling hypothesis" and shifting focus to more efficient models, a paradigm shift from the brute-force approach of the last five years.
Unlike GPUs using slow, dense memory, Cerebras's wafer-sized chip leverages its vast surface area to accommodate faster, less-dense memory. This design sidesteps memory bottlenecks, achieving speeds up to 15 times faster than the fastest GPUs for AI tasks.