Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

As AI agents scale to millions of free users, companies are forced to use cheaper, older, or less capable models to manage costs. This economic pressure is leading to a noticeable decline in quality and a return of hallucinations, creating real-world problems for users like incorrect flight booking information.

Related Insights

Contrary to expectations of falling AI costs, the move from simple chatbots to complex, multi-step agentic systems is causing an explosion in token usage. A single user can trigger hundreds of agents, making expensive frontier models economically unsustainable for many application-layer companies.

While guardrails in prompts are useful, a more effective step to prevent AI agents from hallucinating is careful model selection. For instance, using Google's Gemini models, which are noted to hallucinate less, provides a stronger foundational safety layer than relying solely on prompt engineering with more 'creative' models.

The 'Andy Warhol Coke' era, where everyone could access the best AI for a low price, is over. As inference costs for more powerful models rise, companies are introducing expensive tiered access. This will create significant inequality in who can use frontier AI, with implications for transparency and regulation.

The current subsidized AI subscription model is unsustainable. The inevitable shift to pay-per-token pricing will expose the true cost of inference. For tasks like coding, where AI can "hallucinate" and burn tokens in loops, this creates unpredictable and potentially exorbitant costs, akin to gambling.

The capabilities of free, consumer-grade AI tools are over a year behind the paid, frontier models. Basing your understanding of AI's potential on these limited versions leads to a dangerously inaccurate assessment of the technology's trajectory.

As AI token consumption becomes a major budget item, companies are moving beyond using a single frontier model. Every organization will need a portfolio of models, including cheaper options for less complex tasks, to manage the "madness" of runaway costs.

Microsoft GitHub's dramatic shift to consumption-based pricing for CoPilot, with some model costs increasing 27-fold, is the most direct evidence of the AI industry's unsustainable subsidy model. It reveals the true, previously hidden, compute cost of advanced agentic workflows that companies must now pay.

Granting AI agents autonomy can lead to costly errors. In one experiment, an AI managing a vending machine "hallucinated" a reason to set dynamic prices for protein bars at $15—a 500% margin. It even defended its flawed logic when questioned by its human overseer.

OpenAI's new technique to halve inference costs is being tested on non-paying users, suggesting it likely involves quality compromises. This highlights the universal tension in AI development: optimizing for cost and efficiency almost always comes at the expense of performance, a "no free lunch" reality for developers.

Users notice AI tools getting worse at simple tasks. This may not be a sign of technological regression, but rather a business decision by AI companies to run less powerful, cheaper models to reduce their astronomical operational costs, especially for free-tier users.