Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Fine-tuning, once dismissed as obsolete due to powerful foundation models, is making a comeback. Faced with expensive pay-as-you-go APIs for long-running agents and geopolitical supply risks, companies are returning to fine-tuning smaller, self-hosted models to gain cost control and operational resilience.

Related Insights

For specialized, high-stakes tasks like insurance underwriting, enterprises will favor smaller, on-prem models fine-tuned on proprietary data. These models can be faster, more accurate, and more secure than general-purpose frontier models, creating a lasting market for custom AI solutions.

Instead of relying solely on massive, expensive, general-purpose LLMs, the trend is toward creating smaller, focused models trained on specific business data. These "niche" models are more cost-effective to run, less likely to hallucinate, and far more effective at performing specific, defined tasks for the enterprise.

The "agentic revolution" will be powered by small, specialized models. Businesses and public sector agencies don't need a cloud-based AI that can do 1,000 tasks; they need an on-premise model fine-tuned for 10-20 specific use cases, driven by cost, privacy, and control requirements.

In response to budget blowouts from agentic AI, enterprises are moving beyond simple adoption to active cost management. A new "token efficiency" stack is emerging, featuring tactics like model routing to cheaper alternatives (e.g., DeepSeek) and custom post-trained models to reduce reliance on expensive foundation models.

As enterprises scale AI, the high inference costs of frontier models become prohibitive. The strategic trend is to use large models for novel tasks, then shift 90% of recurring, common workloads to specialized, cost-effective Small Language Models (SLMs). This architectural shift dramatically improves both speed and cost.

An emerging rule from enterprise deployments is to use small, fine-tuned models for well-defined, domain-specific tasks where they excel. Large models should be reserved for generic, open-ended applications with unknown query types where their broad knowledge base is necessary. This hybrid approach optimizes performance and cost.

The push towards enterprise fine-tuning directly challenges the 'bitter lesson'—the theory that massive scale in general models will inevitably outperform specialized, human-curated approaches. The success of this new market segment hinges on proving that customized models can maintain a durable advantage over ever-improving, cheaper generalist models.

Companies like Thinking Machines Lab and Microsoft are shifting the value proposition from raw API access to platforms for enterprise-specific model customization. This addresses corporate needs for data sovereignty, cost control, and specialized performance, creating a new competitive lane focused on enabling customers to own their own models.

Microsoft's strategy lets companies customize proprietary models for specific tasks, achieving near-frontier performance at a fraction of the cost. This 'controlled tuning' approach is a powerful alternative to using expensive general models or relying on potentially inaccessible open-source options from abroad.

To escape platform risk and high API costs, startups are building their own AI models. The strategy involves taking powerful, state-subsidized open-source models from China and fine-tuning them for specific use cases, creating a competitive alternative to relying on APIs from OpenAI or Anthropic.

Enterprises Are Reviving Model Fine-Tuning to Counter Rising API Costs and Geopolitical Risks | RiffOn