We scan new podcasts and send you the top 5 insights daily.
Disruptive innovations (like early cars) typically follow a path of rapid cost reduction (Moore's Law). AI is an exception; despite massive investment, its core compute costs are escalating, not declining, making the classic 'Innovator's Dilemma' analogy flawed.
The tech industry wrongly compares AI to software, which has near-zero marginal costs for new users. In reality, providing access to frontier AI models is a zero-sum game during compute crunches because of immense computational requirements. Servicing another user is expensive, leading to rationed access.
AI software models advance every few months, creating exponential demand. However, the hardware infrastructure like chip fabs operates on two-to-four-year development cycles. This timeline disconnect between software's rapid pace and hardware's slow build-out creates a persistent supply crunch that money alone cannot instantly solve.
Building software traditionally required minimal capital. However, advanced AI development introduces high compute costs, with users reporting spending hundreds on a single project. This trend could re-erect financial barriers to entry in software, making it a capital-intensive endeavor similar to hardware.
A primary risk for major AI infrastructure investments is not just competition, but rapidly falling inference costs. As models become efficient enough to run on cheaper hardware, the economic justification for massive, multi-billion dollar investments in complex, high-end GPU clusters could be undermined, stranding capital.
The cost for a given level of AI capability has decreased by a factor of 100 in just one year. This radical deflation in the price of intelligence requires a complete rethinking of business models and future strategies, as intelligence becomes an abundant, cheap commodity.
In a striking economic anomaly, the cost to rent older NVIDIA H100 AI chips is increasing, not decreasing. This is because the growth in AI's usefulness is outstripping the tripling annual supply of compute. It signals that the value being generated by AI models is growing faster than our ability to manufacture the hardware to run them.
A counterintuitive view of Moore's Law is that for it to hold, the economic value of computation must halve every 18 months because we historically run out of uses for it. The recent rise in H100 GPU rental costs suggests AI is the first application where demand is growing faster than supply, breaking this trend.
Unlike durable infrastructure like railways or fiber optic cables, AI's core component—expensive GPUs—becomes obsolete in just 2-3 years. This creates a permanent, recurring cost, a 'tax on innovation,' making profitability much harder to achieve compared to previous tech revolutions.
Countering the narrative of insurmountable training costs, Jensen Huang argues that architectural, algorithmic, and computing stack innovations are driving down AI costs far faster than Moore's Law. He predicts a billion-fold cost reduction for token generation within a decade.
While hardware gets cheaper (Moore's Law), the competitive pressure to release superior AI models leads to exponentially larger and more complex systems. This results in a higher number of "tokens burned" per query, making the cost of delivering a useful answer actually increase with each new generation.