Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Contrary to the conventional view, the training phase for advanced AI models, especially those using reinforcement learning, now demands more compute resources for inference tasks than for the actual backpropagation process where the model learns.

Related Insights

Unlike simple classification (one pass), generative AI performs recursive inference. Each new token (word, pixel) requires a full pass through the model, turning a single prompt into a series of demanding computations. This makes inference a major, ongoing driver of GPU demand, rivaling training.

Pre-training on internet text data is hitting a wall. The next major advancements will come from reinforcement learning (RL), where models learn by interacting with simulated environments (like games or fake e-commerce sites). This post-training phase is in its infancy but will soon consume the majority of compute.

The compute power required for AI agents to operate ('inference') is a significant new cost. Without an optimized infrastructure to manage these costs, companies risk spending all their AI-driven productivity gains on 'feeding' their digital workers, making the initiative unprofitable.

The pace of AI development is so rapid that a complex inference task assigned to a model could take longer to complete than the time it takes to train and release the next, more powerful version of that same model. This highlights an emerging paradox in the deployment of large-scale AI.

While reinforcement learning (RL) improves model capabilities, it often results in unpredictable, "bursty" computational demands during inference. This complicates serving the model efficiently, as infrastructure must be provisioned for costly peak loads.

RL models can be inefficient during inference. The GPU often sits idle while the CPU calculates rewards, then suddenly gets hit with a massive "burst" of activity. This unpredictable demand makes serving these models costly and complex, requiring conservative GPU allocation.

AI progress was expected to stall in 2024-2025 due to hardware limitations on pre-training scaling laws. However, breakthroughs in post-training techniques like reasoning and test-time compute provided a new vector for improvement, bridging the gap until next-generation chips like NVIDIA's Blackwell arrived.

The traditional separation is disappearing. Fast inference is critical for modern training (e.g., RL rollouts), while training techniques are now essential for inference optimization (e.g., training speculative decoders). This requires engineers to be proficient in both domains.

Previously, the biggest constraint in AI was compute for training next-gen models. Now, the critical bottleneck is providing enough compute for *inference*—the real-time processing of queries from a rapidly growing user base.

AI's computational needs are not just from initial training. They compound exponentially due to post-training (reinforcement learning) and inference (multi-step reasoning), creating a much larger demand profile than previously understood and driving a billion-X increase in compute.

Modern AI Training Now Consumes More Compute for Inference Than Backpropagation | RiffOn