Sohmers argues that the unified left/right opposition to data centers in the West, often based on false information, serves China's strategic interests. As the West slows its infrastructure build-out due to public pressure, China rapidly expands its own data center capacity without similar domestic opposition.
The narrative of AI labs burning cash is misleading. Their core business of selling API access for inference is highly profitable, with margins like Anthropic's reported 80%. Unprofitability stems from the massive, discretionary R&D cost of training next-generation frontier models, not poor unit economics.
While training AI models is a compute-bound problem where more flops yield better results, inference (running the model) is memory-bound. Each token generation requires reading all model weights from memory, making memory bandwidth, not raw processing power, the primary performance bottleneck.
The "memory wall" is a growing chasm between compute power and memory access. In the last decade, GPU flops improved 120-fold, but memory bandwidth only increased 17-fold. This divergence makes memory-bound workloads like AI inference an increasingly severe bottleneck for modern hardware.
The public discourse on slowing down AI development for safety reasons may have a secondary, financial motive. Since training frontier models is the largest expense for AI labs, a collective pause or slowdown serves as a "great way to reduce costs ahead of an IPO," improving financial optics.
Contrary to the view that local models compete with cloud providers, they will likely become powerful drivers of cloud usage. On-device LLMs will constantly process a user's local data and act as agents, deciding when to escalate complex tasks to more powerful cloud models, massively increasing total token volume.
The economics of token pricing are counterintuitive. Processing a cached token is roughly 1/1000th the cost of generating a new one. This means that even when AI providers sell cached tokens at a discount, they make "insane" margins on them, subsidizing the more expensive generation of new tokens.
Overly restrictive AI regulation could concentrate power in the hands of a few large companies, creating a modern "road to serfdom." By making advanced AI development legally exclusive to a small group, technology itself becomes a tool of control, turning the masses into technological "serfs."
With long context windows, the memory for KV caches of user sessions presents a massive scaling challenge. For a 10-trillion parameter model, the collective context for just 50 concurrent users could require more memory (5+ terabytes) than the model weights themselves, flipping the infrastructure priority from model storage to session storage.
Constraints breed innovation. Faced with U.S. export controls on high-end GPUs, Chinese AI labs were forced to develop novel algorithms to work around hardware limitations. This led to breakthroughs like multi-head latent attention (MLA), which reduces memory requirements at the cost of more compute.
The 60x price drop for a million tokens (from $60 to $1) masks a more profound shift in value. A token from today's frontier models is conservatively 100 times more valuable in terms of capability and economic output than a token from five years ago, meaning the value-adjusted cost of intelligence has fallen over 6000x.
