Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

The NeoCloud landscape is expanding beyond pure-play providers. Inference-focused software platforms and custom chip startups are now building their own data centers and buying GPUs. This vertical integration is driven by the need to control supply, meet demand, and improve margins, fundamentally changing their business models.

Related Insights

A new category of "NeoCloud" or "AI-native cloud" is rising, focusing specifically on AI training and inference. Unlike general-purpose clouds like AWS, these platforms are GPU-first, catering to massive AI workloads and addressing the GPU scarcity and different workload patterns found in hyperscalers.

Nebius's talks to acquire AI21 reflect a broader trend where NeoClouds (e.g., CoreWeave) are buying software companies. This strategy aims to create a full-stack platform, offering more than just compute power, thereby increasing customer stickiness and diversifying revenue streams beyond commoditized hardware rentals.

The lines between hardware, cloud, and AI models are blurring. Nvidia is moving up into cloud services, while its customers (hyperscalers) are moving down into custom silicon. This convergence means every major tech company will soon compete across the entire stack, from data centers to APIs.

The AI compute market has stratified into a pyramid. Hyperscalers serve top frontier labs, forcing NeoClouds and inference platforms to build their own data centers. This trickles down, compelling AI startups to seek GPU capacity from an increasingly fragmented landscape, including providers that repurpose crypto mines.

With AI infrastructure spend topping $100B annually, hyperscalers like Amazon and Google are vertically integrating. They now manage everything from data center construction and micro-nuclear power to designing their own custom chips. For them, custom silicon has become a 'rounding error' in their budget and a key strategy to optimize costs.

NVIDIA's new business model involves guaranteeing it will rent back unused GPU capacity from smaller cloud providers. This acts as anchor demand, enabling these 'NeoClouds' to secure financing for massive GPU purchases. It's a strategic move for NVIDIA to build and control its own demand ecosystem, ensuring its chips continue to sell.

General Compute is building its cloud service by intentionally avoiding NVIDIA GPUs. It believes NVIDIA is optimized for low-cost, slow token generation. Instead, it uses ASICs from companies like SambaNova to target the nascent market for high-speed inference (1000+ tokens/sec).

HydroHost's strategy is built on the thesis that data centers are moving beyond being mere cost centers for public clouds. It provides software for them to become "Neo Clouds," serving AI companies directly. This model gives data centers more control and upside, mimicking how crypto miners bypassed clouds for better hardware access.

A new category of cloud providers, "NeoClouds," are built specifically for high-performance GPU workloads. Unlike traditional clouds like AWS, which were retrofitted from a CPU-centric architecture, NeoClouds offer superior performance for AI tasks by design and through direct collaboration with hardware vendors like NVIDIA.

The AI compute crunch isn't only about GPU scarcity. Startups are choosing smaller cloud providers ("neoclouds") over AWS because they offer more flexible terms. They can avoid the large, long-term, and expensive commitments that incumbents often require for high-demand NVIDIA chips.