Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

To navigate millions of public models, Dell built a Model Evaluation Protocol (MEP). This system uses AI agents to test new models across different hardware, plotting performance vs. intelligence on a Pareto curve to provide tailored, data-driven customer recommendations.

Related Insights

Recognizing there is no single "best" LLM, AlphaSense built a system to test and deploy various models for different tasks. This allows them to optimize for performance and even stylistic preferences, using different models for their buy-side finance clients versus their corporate users.

Model leaderboards are misleading. To ensure a consistent user experience, companies must develop their own evaluation suites reflecting their specific workloads. This allows them to swap underlying models for cost or capability reasons with confidence that the customer-facing outcome remains reliable and high-quality.

The primary bottleneck in improving AI is no longer data or compute, but the creation of 'evals'—tests that measure a model's capabilities. These evals act as product requirement documents (PRDs) for researchers, defining what success looks like and guiding the training process.

Leading AI models offer different trade-offs in speed, cost, and capability. A model like GPT-5.6 might be faster and more affordable for 95% of tasks, while a competitor like Fable might be superior for the most complex problems, creating a multi-leader market where different tools are used for different jobs.

The AI model landscape isn't a simple ladder of best to worst. Instead, it's a "spiky" frontier where different models offer unique strengths. For example, one model may excel at complex, niche problems while another is faster, more affordable, and better for collaborative, general-purpose tasks, necessitating a multi-tool approach.

An intelligent AI orchestration layer can achieve a cost-to-accuracy balance superior to any single model. By routing queries to a portfolio of different models (large, small, specialized), it creates a new Pareto frontier, delivering higher success rates at a lower average cost than relying on one "best" model.

To efficiently assess new AI models, develop a personal portfolio of your most critical tasks. This 'reusable evaluation set,' complete with prompts and success criteria, allows you to quickly and consistently benchmark new models against your specific needs, rather than relying on general capabilities.

As AI costs rise, using one powerful frontier model for every task is no longer financially viable. The solution is to create a dedicated "Model Sommelier" role responsible for curating a portfolio of models, continuously testing and selecting the most cost-effective option for each specific business use case.

A significant source of competitive advantage ("alpha") comes from systematically testing various AI models for different tasks. This creates a personal map of which tools are best for specific use cases, ensuring you always use the optimal solution.

The rapid release of new AI models makes it crucial for companies to move beyond industry benchmarks. Developing internal evaluation systems ("evals") is necessary to test and determine which model performs best for unique, high-value business use cases, as model choice is becoming extremely important.