We scan new podcasts and send you the top 5 insights daily.
Counterintuitively, the most advanced models can be cheaper for simple tasks. As a model approaches perfection, it spends fewer tokens on verification steps like running linters, making it more efficient than smaller models that must constantly check their work.
Recent AI model releases are not just cheaper on a per-token basis. They are also engineered to use significantly fewer tokens to generate responses, creating a compound cost-saving effect for users. This signals a strategic shift from raw capability to practical, everyday efficiency.
It's counterintuitive, but using a more expensive, intelligent model like Opus 4.5 can be cheaper than smaller models. Because the smarter model is more efficient and requires fewer interactions to solve a problem, it ends up using fewer tokens overall, offsetting its higher per-token price.
The total cost of an AI task depends on the outcome quality. A high-quality model might use more expensive tokens but achieve the result faster and with fewer attempts, making the overall system cost lower than a cheaper, less effective model.
As AI models become more efficient, cost-per-token is an increasingly misleading metric. A more capable model might be more expensive per token but far cheaper per completed task because it requires fewer steps or revisions. The focus of economic evaluation must shift from the raw input cost to the final output cost.
The binary distinction between "reasoning" and "non-reasoning" models is becoming obsolete. The more critical metric is now "token efficiency"—a model's ability to use more tokens only when a task's difficulty requires it. This dynamic token usage is a key differentiator for cost and performance.
A sophisticated gateway that routes queries to different models based on complexity is key to managing AI costs. Simple tasks go to cheap, open-source models, while difficult ones use the frontier. This "expert pattern" allows token usage to rise while keeping costs flat.
A model with a low per-token price can be more expensive if it's inefficient, verbose, or requires multiple attempts ('overthinking'). The actual invoice depends on the total tokens needed to complete a task, making token efficiency a hidden multiplier that savvy enterprises are now tracking to determine the true cost.
Focusing on token pricing is misleading. A more powerful model may be more expensive per token but significantly cheaper per task because its higher efficiency requires fewer prompts and iterations to achieve a final result. The correct way to measure cost-effectiveness is by the total cost to complete a job, not the price of the raw material.
An AlphaSense study revealed that models with a higher price-per-token, like GPT 5.6 Sol, can complete tasks for a lower total cost than cheaper Chinese models. This is because their superior efficiency requires fewer tokens to achieve a higher-quality result, making simple price comparisons misleading.
Anthropic's Fable 5 costs twice as much per token as its predecessor. However, its increased intelligence leads to fewer errors and more direct solutions, reducing the total tokens needed for a task and making the overall cost more competitive.