Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

LLM agents often resend their entire history with each action, causing context to grow continuously. This pushes them into the least energy-efficient operating zones, where tokens-per-watt halves with each context doubling, making them a worst-case scenario for power consumption.

Related Insights

Contrary to expectations of falling AI costs, the move from simple chatbots to complex, multi-step agentic systems is causing an explosion in token usage. A single user can trigger hundreds of agents, making expensive frontier models economically unsustainable for many application-layer companies.

Agentic AI has a different computational profile than previous generative AI. It uses much larger inputs with high reusability, creating larger KV caches. This change makes GPU architectures, which were optimized for earlier workloads, inefficient for the new demands of agentic inference.

Moving from simple chatbots to autonomous agents creates a massive cost increase. Agents consume 5 to 30 times more tokens because they operate in loops, with each task involving 10-20 separate model calls that carry extensive history, instructions, and tool definitions, rapidly compounding costs.

Simply stuffing all historical data into a large context window is counterproductive. The model's attention gets diluted by repetitive tool logs and intermediate data, making it struggle to find original instructions. This "signal versus noise" problem leads to hallucinations and degraded performance.

Semiconductor design offers a blueprint for AI efficiency. Using tiered models like "voltage islands," adding deterministic gates before LLM calls like "clock gating," and compressing context like "level shifters" can dramatically reduce computational waste and cost.

GPU power consumption remains nearly constant regardless of context size, but throughput collapses as the memory required per token grows. This means longer context windows are exponentially less energy-efficient, measured in tokens per watt. This is called the "1-W law."

Despite massive context windows in new models, AI agents still suffer from a form of 'memory leak' where accuracy degrades and irrelevant information from past interactions bleeds into current tasks. Power users manually delete old conversations to maintain performance, suggesting the issue is a core architectural challenge, not just a matter of context size.

The shift from simple query-based AI to agentic AI, where AI calls itself recursively to solve complex tasks, increases compute demand by orders of magnitude. Most people, especially non-coders, fail to grasp this exponential shift, leading them to consistently underestimate the scale and duration of the AI infrastructure build-out.

The largest driver of future energy consumption for AI won't be human-initiated queries on chatbots. Instead, it will be the massive, continuous "machine-to-machine" traffic generated by autonomous AI agents performing tasks, which will ultimately swamp human-AI interaction and create a runaway demand for compute power.

The simple "tool calling in a loop" model for agents is deceptive. Without managing context, token-heavy tool calls quickly accumulate, leading to high costs ($1-2 per run), hitting context limits, and performance degradation known as "context rot."