Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

To improve reliability of NVIDIA's Vera Rubin systems, CoreWeave developed its own hardware: 'Racky' (a rack manager) and 'Velvi' (a liquid cooling manager). These proprietary tools provide better signals and automation, helping manage the complexity of new, highly integrated GPU racks and speed up deployment.

Related Insights

The AI supply chain is crunched not just by obvious components like TSMC wafers and HBM memory. A significant, often overlooked bottleneck is rack manufacturing—including high-speed cables, connectors, and even sheet metal—which are "sneaky hard" due to extreme power, heat, and signal integrity demands.

As NVIDIA moves to massive rack-scale systems, the primary installation challenge has evolved. It's no longer just about the chips, but the immense cabling and networking connecting them. Diagnosing a single failed cable among kilometers of wiring is now the crucial, non-linear problem, described by one expert as "black magic."

By funding and backstopping CoreWeave, which exclusively uses its GPUs, NVIDIA establishes its hardware as the default for the AI cloud. This gives NVIDIA leverage over major customers like Microsoft and Amazon, who are developing their own chips. It makes switching to proprietary silicon more difficult, creating a competitive moat based on market structure, not just technology.

The Rubin family of chips is sold as a complete "system as a rack," meaning customers can't just swap out old GPUs. This technical requirement creates a forced, expensive upgrade cycle for cloud providers, compelling them to invest heavily in entirely new rack systems to stay competitive.

Despite major tech companies developing their own AI chips, CoreWeave's clients exclusively demand Nvidia hardware. This is attributed to the mature CUDA software platform, which provides an efficient, scalable, and reliable ecosystem that competitors have been unable to replicate.

NVIDIA's complex Blackwell chip transition requires rapid, large-scale deployment to work out bugs. XAI, known for building data centers faster than anyone, serves this role for NVIDIA. This symbiotic relationship helps NVIDIA stabilize its new platform while giving XAI first access to next-generation models.

The new Vera Rubin racks are easier to install not just due to NVIDIA's design improvements, but because customers are now on their "third generation" of deploying rack-scale systems. The difficult rollout of the previous Blackwell chips served as a steep learning curve for data center operators, making them better prepared.

While many focus on physical infrastructure like liquid cooling, CoreWeave's true differentiator is its proprietary software stack. This software manages the entire data center, from power to GPUs, using predictive analytics to gracefully handle component failures and maximize performance for customers' critical AI jobs.

Newer AI cloud providers gain a performance advantage by building their infrastructure entirely on NVIDIA's integrated ecosystem, including specialized networking. Incumbent clouds often must patch their legacy, CPU-centric systems, creating inefficiencies that 'neo-clouds' without technical debt can avoid.

Etched builds its own chips, boards, cold plates, interconnects, and even its own racks. This full-stack ownership allows for extreme parallelization and iteration speed, a key advantage over startups that rely on a fragmented supply chain and multiple vendors.