We scan new podcasts and send you the top 5 insights daily.
The AI compute crunch isn't only about GPU scarcity. Startups are choosing smaller cloud providers ("neoclouds") over AWS because they offer more flexible terms. They can avoid the large, long-term, and expensive commitments that incumbents often require for high-demand NVIDIA chips.
A new category of "NeoCloud" or "AI-native cloud" is rising, focusing specifically on AI training and inference. Unlike general-purpose clouds like AWS, these platforms are GPU-first, catering to massive AI workloads and addressing the GPU scarcity and different workload patterns found in hyperscalers.
Modal Labs provides an infrastructure layer that sits above hyperscalers and specialized AI clouds. Its value is not owning hardware but abstracting the complexity of managing raw GPU capacity. By offering a superior developer experience and a flexible, usage-based model, it solves the variable demand problem inherent in AI applications.
The widely discussed GPU supply crunch is only half the problem. There's a severe shortage of suppliers who can operate data centers with the high reliability and SLAs required for mission-critical inference. Out of many providers, only a handful meet the "gold tier" for operational excellence.
The AI compute market has stratified into a pyramid. Hyperscalers serve top frontier labs, forcing NeoClouds and inference platforms to build their own data centers. This trickles down, compelling AI startups to seek GPU capacity from an increasingly fragmented landscape, including providers that repurpose crypto mines.
A fundamental shift is occurring where startups allocate limited budgets toward specialized AI models and developer tools, rather than defaulting to AWS for all infrastructure. This signals a de-bundling of the traditional cloud stack and a change in platform priorities.
NVIDIA's new business model involves guaranteeing it will rent back unused GPU capacity from smaller cloud providers. This acts as anchor demand, enabling these 'NeoClouds' to secure financing for massive GPU purchases. It's a strategic move for NVIDIA to build and control its own demand ecosystem, ensuring its chips continue to sell.
Once a haven for startups struggling to get GPUs, NeoClouds like CoreWeave have shifted their strategy. They now prioritize serving the largest customers, mirroring the behavior of AWS and Azure and leaving startups with fewer alternative compute options than in 2023.
For leading AI labs like Anthropic and OpenAI, the primary value from cloud partnerships isn't a sales channel but guaranteed access to scarce compute and GPUs. This turns negotiations into a complex, symbiotic bundle covering hardware access, cloud credits, and revenue sharing, where hardware is the most critical component.
A new category of cloud providers, "NeoClouds," are built specifically for high-performance GPU workloads. Unlike traditional clouds like AWS, which were retrofitted from a CPU-centric architecture, NeoClouds offer superior performance for AI tasks by design and through direct collaboration with hardware vendors like NVIDIA.
Newer AI cloud providers gain a performance advantage by building their infrastructure entirely on NVIDIA's integrated ecosystem, including specialized networking. Incumbent clouds often must patch their legacy, CPU-centric systems, creating inefficiencies that 'neo-clouds' without technical debt can avoid.