Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

A 10% daily growth rate means compute needs double weekly. With hardware lead times of several months, a company like Instinct must predict demand far in advance. Buying for projected growth (e.g., 100M users) is a massive capital risk; under-buying means throttling a viral product. This is a unique scaling challenge.

Related Insights

The industry is fixated on the GPU shortage, but the proliferation of AI agents will create massive demand for general-purpose compute, leading to a CPU bottleneck. As millions of agents perform tasks, the availability of CPU cores—not just specialized processors—will become the primary constraint on growth for compute providers.

Unlike human-driven growth, which is limited by population and waking hours, AI agents can operate, replicate, and call each other endlessly. This creates a potentially infinite demand for compute infrastructure, far exceeding previous models and leading to massive, unpredictable strains on providers.

Today's AI computing demand from millions of human users is just the beginning. The real explosion in demand will come from billions of AI "agents" working 24/7. This will double the workforce and create a relentless, round-the-clock need for inference computing, dwarfing current infrastructure requirements.

While GPUs dominate AI hardware discussions, the proliferation of AI agents is causing a significant, often overlooked, CPU shortage. Agents rely on CPUs for web queries, data processing, and other tasks needed to feed GPUs, straining existing infrastructure and driving new demand for companies like Arm and Intel.

The focus in AI has evolved from rapid software capability gains to the physical constraints of its adoption. The demand for compute power is expected to significantly outstrip supply, making infrastructure—not algorithms—the defining bottleneck for future growth.

Unlike traditional software, the success of an AI company is inextricably linked to securing and planning for GPU capacity. This has elevated compute strategy and demand forecasting to a critical, board-level CEO responsibility. Misjudging this can be fatal, as some labs have been caught off guard by unexpected growth.

Unlike on-demand tools like code generators, proactive agents like Instinct are constantly working in the background—waking up, scanning information, and anticipating user needs. This "always-on" nature means compute consumption is not tied to direct user interaction, creating a demand that is orders of magnitude larger than previous AI products.

Beyond the well-known GPU scarcity for training AI, a massive CPU shortage looms for deployment. Providing a single AI agent for every knowledge worker would require 40 times the current annual global CPU production, highlighting a critical, under-discussed infrastructure challenge for the agentic web.

Previously, the biggest constraint in AI was compute for training next-gen models. Now, the critical bottleneck is providing enough compute for *inference*—the real-time processing of queries from a rapidly growing user base.

The transition from chatbots to autonomous 'agentic' AI represents a fundamental step-change. These agents, which execute complex tasks independently, have already increased the demand for computational power by 1000x, creating a massive, ongoing need for new infrastructure and hardware.

Viral AI Agents Face an Unprecedented Compute Procurement Dilemma | RiffOn