Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

The market incorrectly feared that the highly capable Kimi model would create a compute glut. In reality, it's a massive 2.8 trillion-parameter model that requires huge, well-networked GPU superclusters to run effectively, thereby increasing demand for NVIDIA's high-end hardware.

Related Insights

NVIDIA's revenue growth is speeding up even as its revenue base expands massively, a rare feat that defies the "law of large numbers." This suggests strong network effects and a dominant market position are creating a self-reinforcing cycle of demand for its AI hardware.

World models are algorithmically more intense than language models, pushing computation (flops) much harder relative to memory access. This unique computational pattern will create a market for specialized chips optimized specifically for these workloads, leading to a divergence from the current hardware landscape built for LLMs.

Hardware shortages act as a catalyst for software innovation. The 'Kimi moment,' where a Chinese model introduced major memory efficiency improvements, demonstrates a recurring pattern: when a component like memory becomes a bottleneck, the ecosystem responds with algorithmic breakthroughs to reduce demand for it.

When an efficient model like DeepSeek was released, Nebius's stock fell on fears of reduced compute demand. Internally, they had their best sales week ever. Cheaper intelligence makes new products economically viable, increasing overall compute consumption, not decreasing it.

Contrary to fears that efficient models hurt NVIDIA, large open-source models like Kimi K3 (2.8T+ parameters) are a net positive. Their sheer size necessitates large-scale GPU clusters for inference just to store the weights, driving demand for high-end, scale-up hardware like NVIDIA's NVL72 regardless of algorithmic efficiency.

The "CUDA moat" is misunderstood. NVIDIA's true advantage is that major open-source models (e.g., from DeepSeek, Alibaba) are co-designed for its GPUs. This creates a powerful downstream effect where developers must use NVIDIA hardware to run the best available models, regardless of the programming layer.

A critical, under-discussed constraint on Chinese AI progress is the compute bottleneck caused by inference. Their massive user base consumes available GPU capacity serving requests, leaving little compute for the R&D and training needed to innovate and improve their models.

Despite developing frontier-level AI models like Kimi K3, Chinese labs are severely compute-constrained. The Kimi K3 launch quickly overwhelmed servers, revealing a lack of GPU infrastructure for large-scale inference. This shifts the US-China competition focus from model benchmarks to industrial capacity and data center dominance.

The value unlocked by frontier AI models is expanding so rapidly that there isn't enough hardware to meet demand. This scarcity ensures that not just the top lab (like OpenAI), but also second and third-tier competitors, will operate at full capacity with strong margins.

AI's computational needs are not just from initial training. They compound exponentially due to post-training (reinforcement learning) and inference (multi-step reasoning), creating a much larger demand profile than previously understood and driving a billion-X increase in compute.