Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

The economics of token pricing are counterintuitive. Processing a cached token is roughly 1/1000th the cost of generating a new one. This means that even when AI providers sell cached tokens at a discount, they make "insane" margins on them, subsidizing the more expensive generation of new tokens.

Related Insights

While the cost-per-token is decreasing as models become more efficient, this efficiency gain drives a massive increase in new use cases and overall consumption. This economic principle, Jevons Paradox, explains why total enterprise spending on model inference is skyrocketing, even as the unit cost falls.

Contrary to the narrative of burning cash, major AI labs are likely highly profitable on the marginal cost of inference. Their massive reported losses stem from huge capital expenditures on training runs and R&D. This financial structure is more akin to an industrial manufacturer than a traditional software company, with high upfront costs and profitable unit economics.

The narrative of AI labs burning cash is misleading. Their core business of selling API access for inference is highly profitable, with margins like Anthropic's reported 80%. Unprofitability stems from the massive, discretionary R&D cost of training next-generation frontier models, not poor unit economics.

As AI models become more efficient, cost-per-token is an increasingly misleading metric. A more capable model might be more expensive per token but far cheaper per completed task because it requires fewer steps or revisions. The focus of economic evaluation must shift from the raw input cost to the final output cost.

A model with a low per-token price can be more expensive if it's inefficient, verbose, or requires multiple attempts ('overthinking'). The actual invoice depends on the total tokens needed to complete a task, making token efficiency a hidden multiplier that savvy enterprises are now tracking to determine the true cost.

The Alkin-Allen effect suggests that as a fixed cost (expensive compute) dominates, users will pay a large premium for the highest quality good (the most efficient model). This allows top labs to charge much higher margins for models that economize on costly compute by using fewer tokens to achieve the same result.

The business model for foundation models could become incredibly lucrative if providers can subtly adjust the "dials"—like token cost or consumption per task—to manage profitability. This creates an opaque market where they extract enormous margins, unless open competition forces transparency and commoditization.

Focusing on token pricing is misleading. A more powerful model may be more expensive per token but significantly cheaper per task because its higher efficiency requires fewer prompts and iterations to achieve a final result. The correct way to measure cost-effectiveness is by the total cost to complete a job, not the price of the raw material.

AI vendors can increase effective costs without altering their price-per-token sheets. When Anthropic updated its model, a new tokenizer produced ~30% more tokens for the same text. While the sticker price was unchanged, real-world bills grew 12-27%—a classic case of 'shrinkflation' applied to AI.

An AlphaSense study revealed that models with a higher price-per-token, like GPT 5.6 Sol, can complete tasks for a lower total cost than cheaper Chinese models. This is because their superior efficiency requires fewer tokens to achieve a higher-quality result, making simple price comparisons misleading.