Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

NeoClouds that use non-standard networking hardware instead of the NVIDIA default face a crucial trade-off. New open-source AI software is almost exclusively developed and optimized for NVIDIA's network. This means customers on custom networks can wait months for support for new features, a costly delay in the fast-moving AI space.

Related Insights

A new category of "NeoCloud" or "AI-native cloud" is rising, focusing specifically on AI training and inference. Unlike general-purpose clouds like AWS, these platforms are GPU-first, catering to massive AI workloads and addressing the GPU scarcity and different workload patterns found in hyperscalers.

As chip manufacturers like NVIDIA release new hardware, inference providers like Base10 absorb the complexity and engineering effort required to optimize AI models for the new chips. This service is a key value proposition, saving customers from the challenging process of re-optimizing workloads for new hardware.

Emerging cloud providers (“NeoClouds”) are sticking exclusively with NVIDIA, despite alternatives from AMD. The perceived performance risk is too high, as customers demand state-of-the-art inference speed and providers can't risk a multi-billion dollar investment on a non-NVIDIA stack that might offer lower throughput.

Despite major tech companies developing their own AI chips, CoreWeave's clients exclusively demand Nvidia hardware. This is attributed to the mature CUDA software platform, which provides an efficient, scalable, and reliable ecosystem that competitors have been unable to replicate.

NVIDIA is evolving from a pure hardware provider to an integrated AI platform company with its Nemo Switchyard model router. This software offering creates stickiness for its hardware stack, providing a compelling, all-in-one solution for enterprises that operate their own data centers and want to optimize AI workflows.

Specialized AI clouds (NeoClouds) like CoreWeave emerged because hyperscalers' strengths—such as custom networking and security for multi-tenancy—were detrimental to the performance of large-scale, single-tenant AI workloads. This performance gap created a significant market opening.

General Compute is building its cloud service by intentionally avoiding NVIDIA GPUs. It believes NVIDIA is optimized for low-cost, slow token generation. Instead, it uses ASICs from companies like SambaNova to target the nascent market for high-speed inference (1000+ tokens/sec).

A new category of cloud providers, "NeoClouds," are built specifically for high-performance GPU workloads. Unlike traditional clouds like AWS, which were retrofitted from a CPU-centric architecture, NeoClouds offer superior performance for AI tasks by design and through direct collaboration with hardware vendors like NVIDIA.

Newer AI cloud providers gain a performance advantage by building their infrastructure entirely on NVIDIA's integrated ecosystem, including specialized networking. Incumbent clouds often must patch their legacy, CPU-centric systems, creating inefficiencies that 'neo-clouds' without technical debt can avoid.

Nvidia is developing networking technology that allows non-Nvidia AI chips to work together. This strategic move ensures customers remain within Nvidia's ecosystem, even if they don't buy Nvidia's GPUs, by capturing them at the crucial interconnect layer.