Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Research shows that for agentic tasks, accuracy often peaks at an intermediate cost and then saturates or declines. Spending more tokens not only fails to yield better results but also comes with high variance and unpredictable costs, as models cannot accurately forecast their own consumption.

Related Insights

Contrary to expectations of falling AI costs, the move from simple chatbots to complex, multi-step agentic systems is causing an explosion in token usage. A single user can trigger hundreds of agents, making expensive frontier models economically unsustainable for many application-layer companies.

Incentivizing high AI token usage is not waste, but a form of R&D. In the new agentic paradigm, there are no best practices. Mass experimentation, even with failures, is the only way to discover future workflows and avoid being left behind.

Moving from simple chatbots to autonomous agents creates a massive cost increase. Agents consume 5 to 30 times more tokens because they operate in loops, with each task involving 10-20 separate model calls that carry extensive history, instructions, and tool definitions, rapidly compounding costs.

Progress in complex, long-running agentic tasks is better measured by tokens consumed rather than raw time. Improving token efficiency, as seen from GPT-5 to 5.1, directly enables more tool calls and actions within a feasible operational budget, unlocking greater capabilities.

Track the number of tokens each autonomous coding task consumes. Unexpectedly high token usage signals that your agent encountered problems, highlighting opportunities to improve its tooling, instructions, or environmental checks for future efficiency gains.

While early generative AI costs were negligible, the shift to complex, multi-step agentic workflows is causing a massive spike in token usage. This has elevated cost optimization and ROI from a minor concern to a C-suite priority for the first time.

The total cost of an AI task depends on the outcome quality. A high-quality model might use more expensive tokens but achieve the result faster and with fewer attempts, making the overall system cost lower than a cheaper, less effective model.

As AI models become more efficient, cost-per-token is an increasingly misleading metric. A more capable model might be more expensive per token but far cheaper per completed task because it requires fewer steps or revisions. The focus of economic evaluation must shift from the raw input cost to the final output cost.

In complex, multi-step tasks, overall cost is determined by tokens per turn and the total number of turns. A more intelligent, expensive model can be cheaper overall if it solves a problem in two turns, while a cheaper model might take ten turns, accumulating higher total costs. Future benchmarks must measure this turn efficiency.

An AI model might have a low cost per token but be 'token hungry,' requiring more tokens to complete a task. This makes it more expensive overall than a model with a higher per-token cost but greater efficiency. Evaluating models on a 'cost per task' basis provides a more accurate ROI.

More Token Usage in AI Agents Fails to Improve Accuracy | RiffOn