Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

The market wrongly feared that the Chinese model Kimmy would reduce compute needs. In reality, its massive size (2.8T parameters) and high capability create new use cases for agentic AI, fueling demand for large NVIDIA-powered GPU clusters, not diminishing it.

Related Insights

The market incorrectly feared that the highly capable Kimi model would create a compute glut. In reality, it's a massive 2.8 trillion-parameter model that requires huge, well-networked GPU superclusters to run effectively, thereby increasing demand for NVIDIA's high-end hardware.

Hardware shortages act as a catalyst for software innovation. The 'Kimi moment,' where a Chinese model introduced major memory efficiency improvements, demonstrates a recurring pattern: when a component like memory becomes a bottleneck, the ecosystem responds with algorithmic breakthroughs to reduce demand for it.

When an efficient model like DeepSeek was released, Nebius's stock fell on fears of reduced compute demand. Internally, they had their best sales week ever. Cheaper intelligence makes new products economically viable, increasing overall compute consumption, not decreasing it.

Contrary to fears that efficient models hurt NVIDIA, large open-source models like Kimi K3 (2.8T+ parameters) are a net positive. Their sheer size necessitates large-scale GPU clusters for inference just to store the weights, driving demand for high-end, scale-up hardware like NVIDIA's NVL72 regardless of algorithmic efficiency.

The "CUDA moat" is misunderstood. NVIDIA's true advantage is that major open-source models (e.g., from DeepSeek, Alibaba) are co-designed for its GPUs. This creates a powerful downstream effect where developers must use NVIDIA hardware to run the best available models, regardless of the programming layer.

The rise of agentic AI and reinforcement learning is increasing the need for powerful CPUs located near GPUs. Cloud provider Nebius notes CPU requirements can be a high multiple of the GPU count, fueling a new demand cycle.

The current compute crunch isn't just a supply issue. It's because new AI models are so much more capable that they unlock a total addressable market (TAM) of valuable tasks that grows exponentially, far outpacing the linear or geometric growth of compute supply.

The value unlocked by frontier AI models is expanding so rapidly that there isn't enough hardware to meet demand. This scarcity ensures that not just the top lab (like OpenAI), but also second and third-tier competitors, will operate at full capacity with strong margins.

The market isn't a battle between proprietary frontier models and open-source alternatives. Instead, both are seeing parabolic growth. While open-source becomes more capable for simple tasks, the demand for cutting-edge capabilities unlocked by frontier models is also expanding rapidly, creating a positive-sum environment.

The success of personal AI assistants signals a massive shift in compute usage. While training models is resource-intensive, the next 10x in demand will come from widespread, continuous inference as millions of users run these agents. This effectively means consumers are buying fractions of datacenter GPUs like the GB200.