GPU power consumption remains nearly constant regardless of context size, but throughput collapses as the memory required per token grows. This means longer context windows are exponentially less energy-efficient, measured in tokens per watt. This is called the "1-W law."
LLM agents often resend their entire history with each action, causing context to grow continuously. This pushes them into the least energy-efficient operating zones, where tokens-per-watt halves with each context doubling, making them a worst-case scenario for power consumption.
Semiconductor design offers a blueprint for AI efficiency. Using tiered models like "voltage islands," adding deterministic gates before LLM calls like "clock gating," and compressing context like "level shifters" can dramatically reduce computational waste and cost.
Research shows that for agentic tasks, accuracy often peaks at an intermediate cost and then saturates or declines. Spending more tokens not only fails to yield better results but also comes with high variance and unpredictable costs, as models cannot accurately forecast their own consumption.
Measuring 'tokens per watt' is insufficient. The key metric is 'tokens per completed, accepted task per watt.' This reframes efficiency as a yield problem, where tokens spent on rejected or useless outputs are like defective silicon, wasting both compute and power.
