Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

The emergence of specialized models like JEV signals a shift away from a "one model fits all" approach. Instead of forcing a single, expensive LLM to perform all tasks, companies will build complex architectures using a "model stack." This involves using fast judgment models for routing and then invoking generative models only when necessary.

Related Insights

Relying on a single frontier model is risky and inefficient. The next phase of AI will involve intelligently routing queries to the most appropriate model—be it cheaper, faster, or local. This will redistribute value from a few dominant labs to a long tail of specialized models, maturing the ecosystem.

The future of enterprise AI isn't choosing one provider. Instead, companies will use a "composable model" approach, routing queries to a combination of powerful frontier models and their own fine-tuned open-source models. This strategy, dubbed the "council of LLMs," optimizes for cost, performance, and specialization on proprietary data.

According to Meta's CTO, the era of one monolithic model doing everything is over. The current frontier involves using a 'harness' that intelligently routes tasks to a collection of different, specialized models based on cost, latency, and capability.

Just as developers use various databases for different needs, AI applications will rely on a "constellation" of specialized models. Some tasks will require expensive, high-reasoning models, while others will prioritize low-latency or low-cost models. The market will become heterogeneous, not monolithic.

Instead of relying on one powerful model for all tasks, the leading strategy is 'smart routing'—using a panel of models and directing each task to the most appropriate one. This compound architecture demonstrably beats single frontier models on both cost and performance.

Enterprises will shift from relying on a single large language model to using orchestration platforms. These platforms will allow them to 'hot swap' various models—including smaller, specialized ones—for different tasks within a single system, optimizing for performance, cost, and use case without being locked into one provider.

Instead of relying on a single large AI model, companies are adopting "model orchestration" to control costs. This involves using a router to send prompts to the most appropriate model based on the task, often cascading from cheap, small models to more expensive ones only when necessary.

Breakthroughs will emerge from 'systems' of AI—chaining together multiple specialized models to perform complex tasks. GPT-4 is rumored to be a 'mixture of experts,' and companies like Wonder Dynamics combine different models for tasks like character rigging and lighting to achieve superior results.

Companies no longer chase the single most powerful AI model. The new standard is creating a sophisticated architecture of multiple models, matching the right tool to the right task based on capability, efficiency, and cost, which allows for greater optimization across the enterprise.

Rather than relying on one powerful model, sophisticated users are creating workflows that delegate tasks to different models based on capability and cost. This makes the 'division of labor'—how models like Fable, Opus, and Sonnet are orchestrated—the key strategic unit for building efficient AI systems.

Businesses Will Assemble AI "Model Stacks" Instead of Relying on a Single LLM | RiffOn