Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

The mean time between failure for GPUs depends heavily on their use. Chips used for intensive model training have much higher failure rates than those used for inference. Investors often conflate the two, underestimating the high churn and replacement costs for training hardware.

Related Insights

The relentless pace of new AI models, which perform best on the latest hardware, drastically shortens the effective lifespan of GPUs. This changes the traditional 6-year depreciation model and complicates the financial calculus for building data centers versus renting cloud capacity.

The initial deployment of a new AI cluster sees a high failure rate, with 10-15% of new-generation GPUs like Blackwell needing to be returned or reseated. This "infant mortality" is a standard operational challenge for data centers, underscoring the physical difficulties of scaling AI infrastructure with bleeding-edge chips.

Separating inference into "prefill" (memory-bound) and "decode" (bandwidth-bound) tasks is a game-changer for hardware longevity. It allows older GPUs to be used for prefill tasks indefinitely, extending their useful economic life from 3-4 years to 10-15 years, a boon for data centers and their financiers.

Big tech companies are accounting for AI data centers over a 25-year lifespan. However, the core components, like GPUs, have a much shorter 2-3 year innovation cycle. This discrepancy creates a significant financial risk, as companies could be left with billions in overvalued, obsolete assets on their books.

While the industry standard is a six-year depreciation for data center hardware, analyst Dylan Patel warns this is risky for GPUs. Rapid annual performance gains from new models could render older chips economically useless long before they physically fail.

Contrary to fears of rapid obsolescence, new domain-specific accelerators (DSAs) can be paired with older GPUs to handle specific tasks. This disaggregated approach extends the useful life of GPUs to 10-15 years, lowering financing costs for compute providers and invalidating bear cases.

Countering the narrative of rapid burnout, CoreWeave cites historical data showing a nearly 10-year service life for older NVIDIA GPUs (K80) in major clouds. Older chips remain valuable for less intensive tasks, creating a tiered system where new chips handle frontier models and older ones serve established workloads.

Unlike durable infrastructure like railways or fiber optic cables, AI's core component—expensive GPUs—becomes obsolete in just 2-3 years. This creates a permanent, recurring cost, a 'tax on innovation,' making profitability much harder to achieve compared to previous tech revolutions.

Responding to the AI bubble concern, IBM's CEO notes high GPU failure rates are a design choice for performance. Unlike sunken costs from past bubbles, these "stranded" hardware assets can be detuned to run at lower power, increasing their resilience and extending their useful life for other tasks.

Accusations that hyperscalers "cook the books" by extending GPU depreciation misunderstand hardware lifecycles. Older chips remain at full utilization for less demanding tasks. High operational costs (power, cooling) provide a natural economic incentive to retire genuinely unprofitable hardware, invalidating claims of artificial earnings boosts.