Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

There exists a "scaling law for bandwidth." By solving the memory bottleneck, chip designers enable new AI model architectures that are computationally cheaper. For example, making Mixture-of-Experts (MoE) models significantly sparser saves flops but is prohibitive on today's bandwidth-starved GPUs. High-bandwidth chips make these more efficient architectures viable.

Related Insights

Existing AI chips force a trade-off: high-throughput HBM memory (NVIDIA, Google) has high latency, while low-latency SRAM memory (Grok) has poor throughput. MatX's architecture combines both, putting model weights in fast SRAM and inference data in high-capacity HBM to achieve both low latency and high throughput.

The "memory wall" is a growing chasm between compute power and memory access. In the last decade, GPU flops improved 120-fold, but memory bandwidth only increased 17-fold. This divergence makes memory-bound workloads like AI inference an increasingly severe bottleneck for modern hardware.

NVIDIA is reportedly considering releasing its next-gen Rubin GPUs with less memory than announced due to supply constraints on high-bandwidth memory (HBM). This suggests fundamental hardware limitations, not just algorithms or data, may soon become the primary bottleneck slowing the pace of AI model growth.

AI workloads are limited by memory bandwidth, not capacity. While commodity DRAM offers more bits per wafer, its bandwidth is over an order of magnitude lower than specialized HBM. This speed difference would starve the GPU's compute cores, making the extra capacity useless and creating a massive performance bottleneck.

As AI models evolve to mirror the human brain, their memory requirements are skyrocketing, creating a 'RAMpocalypse.' The industry's focus will shift from being purely compute-centric to a dual focus on memory and compute, making high-bandwidth memory a critical and scarce resource.

A key trend in AI models is "dynamism"—the ability to vary computation and memory usage per token, as seen in Mixture-of-Experts (MoE) architectures. Current hardware, designed before this trend, is inefficient. New chips must be built to accelerate these dynamic computations.

While many focus on compute metrics like FLOPS, the primary bottleneck for large AI models is memory bandwidth—the speed of loading weights into the GPU. This single metric is a better indicator of real-world performance from one GPU generation to the next than raw compute power.

For 20 years, chip design focused on scaling flops (a million-fold increase), while memory bandwidth saw only a 40x increase, creating a massive bottleneck for large AI models. Fractile's core bet is that prioritizing extremely high-bandwidth memory access is the key to unlocking the next level of AI performance and efficiency.

Unlike GPUs using slow, dense memory, Cerebras's wafer-sized chip leverages its vast surface area to accommodate faster, less-dense memory. This design sidesteps memory bottlenecks, achieving speeds up to 15 times faster than the fastest GPUs for AI tasks.

Google's AI chips are designed with unusually high-bandwidth memory. This architecture suggests Google is preparing for a future of continuously learning AI that self-improves 24/7, moving beyond the current cadence of periodic, static model releases.