Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Top AI companies like Meta, Microsoft, and OpenAI are so desperate for compute that they willingly manage systems from both NVIDIA and AMD. This urgent need for capacity overrides the significant operational complexity of writing software that works across different hardware vendors.

Related Insights

Firms like OpenAI and Meta claim a compute shortage while also exploring selling compute capacity. This isn't a contradiction but a strategic evolution. They are buying all available supply to secure their own needs and then arbitraging the excess, effectively becoming smaller-scale cloud providers for AI.

The widely discussed GPU supply crunch is only half the problem. There's a severe shortage of suppliers who can operate data centers with the high reliability and SLAs required for mission-critical inference. Out of many providers, only a handful meet the "gold tier" for operational excellence.

Meta is deprioritizing its custom silicon program, opting for large orders of AMD's chips. This reflects a broader trend among hyperscalers: the urgent need for massive, immediate compute power is outweighing the long-term strategic goal of self-sufficiency and avoiding the "Nvidia tax."

Meta's massive, multi-billion dollar deal for millions of Nvidia GPUs signifies a strategic pivot. After pursuing custom silicon and AMD partnerships to avoid the 'Nvidia tax,' Meta is now committing to Nvidia for the foreseeable future. This move aims to secure a dominant supply of leading AI chips at world-leading scale, prioritizing performance and availability over cost diversification.

Anthropic's strategy of running workloads on diverse chips (NVIDIA, Google TPU, AWS Trainium) is less about long-term diversification and more about immediate survival. In a market where compute is severely constrained, the ability to utilize any available chip becomes a critical competitive advantage, forcing deep technical competence across architectures.

To meet surging demand, Anthropic is diversifying its chip supply beyond NVIDIA. An early adopter of Google's TPUs and Amazon's Tranium, its exploration of Microsoft's custom chips reflects a core philosophy of leveraging any available compute resource rather than committing to a single architecture.

The demand for AI processing power so vastly outstrips supply that it creates a "compute deficit." This forces major AI players to adopt any viable chip solution they can find, including from AMD. It's not about being better than NVIDIA; it's about being available, ensuring a market for second and third-tier suppliers.

Mark, CTO of AMD, states that the explosion of agentic AI workflows has created an unforeseen demand for a balanced compute architecture. These complex, multi-step processes require a CPU to GPU ratio approaching 1:1, a significant shift from traditional GPU-heavy AI training and inference models.

The inference market is too large to remain monolithic. It will fragment into specialized platforms for different use cases like real-time video, long-running agents, or language models. This specialization will extend to hardware, with high-throughput, low-latency-need tasks (like agents) favoring cheaper AMD/Intel chips over NVIDIA's top GPUs.

Previously, the bottleneck for AI labs was researcher time, making Nvidia's easy-to-use CUDA ecosystem dominant. Now, the biggest cost is compute capacity itself, creating massive economic incentives for labs to adopt cheaper, even if less convenient, competing chips from AMD or Google.