Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

A Databricks study found that for coding tasks, the cheaper Sonnet model was ultimately more expensive per task ($2.09) than the premium Opus model ($1.94). The cheaper model required more iterations and reasoning to achieve the same result, proving that the lowest token price doesn't guarantee the lowest total cost.

Related Insights

It's counterintuitive, but using a more expensive, intelligent model like Opus 4.5 can be cheaper than smaller models. Because the smarter model is more efficient and requires fewer interactions to solve a problem, it ends up using fewer tokens overall, offsetting its higher per-token price.

When evaluating AI agents, the total cost of task completion is what matters. A model with a higher per-token cost can be more economical if it resolves a user's query in fewer turns than a cheaper, less capable model. This makes "number of turns" a primary efficiency metric.

Newer AI models may have low per-token prices but are often "token hungry," requiring more tokens to complete a task. This can make them more expensive overall. The true measure of economic viability is the final cost-per-task, not the misleading per-token price.

While models like Grok 4.5 are significantly cheaper per task, their speed enables users to complete work 10-15x faster. This doesn't result in cost savings; instead, users fill the extra time with more tasks, dramatically increasing output and overall token consumption.

OpenAI's GPT-5.5 is more expensive per token, but a new evaluation framework is emerging. The key metric isn't raw cost, but the model's efficiency in solving a problem. This 'intelligence per dollar' reframes cost analysis around performance and compute, where more expensive models can be cheaper overall if they solve tasks more efficiently.

Evaluating AI models on cost-per-token is misleading because it ignores the hidden cost of human labor to fix failures. The true 'cost per successful task' is a business metric that accounts for both the API invoice and the payroll expense for rework, revealing a more accurate total cost of ownership.

A model with a low per-token price can be more expensive if it's inefficient, verbose, or requires multiple attempts ('overthinking'). The actual invoice depends on the total tokens needed to complete a task, making token efficiency a hidden multiplier that savvy enterprises are now tracking to determine the true cost.

A significant, hidden component of AI costs comes from 'reasoning tokens'—the model's internal thinking process before generating an answer. This is invisible to the user but billed at the higher output-token rate and can increase a request's cost by 4x to 20x, similar to a restaurant charging for 'kitchen time'.

Sticker price per token is a misleading metric for AI models. A cheaper model may require more retries or reasoning, making it more expensive overall. The true metric is 'cost per accepted task,' which accounts for total resources needed to get a reliable, usable result, providing a true apples-to-apples comparison.

An AI model might have a low cost per token but be 'token hungry,' requiring more tokens to complete a task. This makes it more expensive overall than a model with a higher per-token cost but greater efficiency. Evaluating models on a 'cost per task' basis provides a more accurate ROI.