Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Despite narratives of scarcity, GPU utilization in some data centers is as low as 35-40%. This is due to colossal hoarding and double-ordering by companies terrified of being caught short in a future parabolic moment. This behavior creates a phantom demand and points to a future supply glut.

Related Insights

Firms like OpenAI and Meta claim a compute shortage while also exploring selling compute capacity. This isn't a contradiction but a strategic evolution. They are buying all available supply to secure their own needs and then arbitraging the excess, effectively becoming smaller-scale cloud providers for AI.

Major AI labs plan and purchase GPUs on multi-year timelines. This means NVIDIA's current stellar earnings reports reflect long-term capital commitments, not necessarily current consumer usage, potentially masking a slowdown in services like ChatGPT.

The advertised per-hour GPU cost is misleading. Because research workloads are spiky and unpredictable, labs over-provision compute. This rampant underutilization means the effective price paid is often 10 times higher than the marketed rate, creating massive deadweight loss.

Large tech companies are buying up compute from smaller cloud providers not for immediate need, but as a defensive strategy. By hoarding scarce GPU capacity, they prevent competitors from accessing critical resources, effectively cornering the market and stifling innovation from rivals.

The current GPU shortage is a temporary state. In any commodity-like market, a shortage creates a glut, and vice-versa. The immense profits generated by companies like NVIDIA are a "bat signal" for competition, ensuring massive future build-out and a subsequent drop in unit costs.

The availability of compute from Meta and XAI doesn't indicate a market-wide surplus. Instead, it points to a compute allocation problem. Massive capacity is concentrated in the hands of companies that currently lack sufficient internal inference demand for their own models, while other parts of the market remain constrained.

Contrary to expectations of easing supply, the GPU shortage has intensified since 2023. With clearer AI business models, mega-customers like OpenAI and Anthropic are spending even more aggressively, creating a fierce bidding war that pushes startups out.

The narrative of insatiable AI compute demand is partially a bubble. It's fueled by inefficient early models ("token maxing") and a culture where tech executives brag about their AI spending as a status symbol, a behavior not seen with traditional cloud costs. This suggests demand could normalize.

To avoid losing their allocated GPUs, some AI researchers are "gaming the system" by running repetitive, useless tasks to create the illusion of high utilization. This behavior stems from intense internal competition for scarce computing resources, leading to inefficient practices designed to protect individual access to hardware.

A major paradox exists in AI development: companies are desperate for scarce GPUs, yet often fail to use them efficiently. Even well-funded labs like XAI report model flops utilization as low as 11%, far below the 40% practical target, due to inconsistent workloads and data transfer bottlenecks.

Massive Hoarding of GPUs by Tech Giants Is Masking Low Utilization Rates | RiffOn