Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Hyperscalers like AWS charge 2-3 times more for identical GPUs compared to smaller clouds. This premium isn't for the chip itself but for a bundled product including a legacy of software, compliance, safety, and long-standing enterprise relationships that create customer stickiness.

Related Insights

Nvidia's staggering revenue growth and 56% net profit margins are a direct cost to its largest customers (AWS, Google, OpenAI). This incentivizes them to form a defacto alliance to develop and adopt alternative chips to commoditize the accelerator market and reclaim those profits.

Tech giants often initiate custom chip projects not with the primary goal of mass deployment, but to create negotiating power against incumbents like NVIDIA. The threat of a viable alternative is enough to secure better pricing and allocation, making the R&D cost a strategic investment.

Providing GPUs-as-a-Service is not a durable business because customers can easily switch providers. The key to customer retention and high net dollar retention (NDR) is the software layer built on top of the hardware. This software, which handles the complexities of inference, creates the actual stickiness.

For leading AI labs like Anthropic and OpenAI, the primary value from cloud partnerships isn't a sales channel but guaranteed access to scarce compute and GPUs. This turns negotiations into a complex, symbiotic bundle covering hardware access, cloud credits, and revenue sharing, where hardware is the most critical component.

Accessing next-generation GPUs at scale is no longer a simple purchase. The market now demands three-to-five-year commitments with a significant portion (20-30%) of the total contract value paid upfront. This makes a company's cost of capital a critical competitive factor in acquiring compute capacity.

A liquid futures market for GPU compute would create price transparency, threatening the business models of hyperscale cloud providers. These giants benefit from opaque, bundled pricing and controlling supply. They will naturally resist the standardization and transparency that an open futures market would bring.

Despite predictions of commoditization, the AI inference layer remains competitive. The market is supply-constrained, and GPU makers like NVIDIA intentionally avoid customer concentration with hyperscalers, creating space for specialized, innovative providers to thrive.

In a power-constrained world, total cost of ownership is dominated by the revenue a data center can generate per watt. A superior NVIDIA system producing multiples more revenue makes the hardware cost irrelevant. A competitor's chip would be rejected even if free due to the high opportunity cost.

Major AI labs aren't just evaluating Google's TPUs for technical merit; they are using the mere threat of adopting a viable alternative to extract significant concessions from Nvidia. This strategic leverage forces Nvidia to offer better pricing, priority access, or other favorable terms to maintain its market dominance.

When all cloud providers offer the same NVIDIA hardware, they are forced to compete on price, eroding margins. By integrating specialized hardware like SambaNova's, they can offer premium, differentiated services—such as faster inference on larger models—allowing them to charge more and improve overall business economics.