We scan new podcasts and send you the top 5 insights daily.
To evaluate AI for commerce, Alibaba developed a custom benchmark using 107 real business tasks. It goes beyond simple accuracy, measuring agents on pass rate, completion time, and crucially, cost. This ensures the solution is not just effective but also affordable and useful for small businesses.
Alibaba.com defines a successful AI agent as the product of three essential components: the core intelligence (Model), the proprietary logic and tools (Harness), and real-world business data (Context). The multiplicative relationship means all three are critical; if one is weak, the entire agent fails.
Many influential AI model benchmarks focus on raw capabilities, like problem-solving accuracy, but neglect a critical business metric: the cost to achieve that result. Future benchmarks must incorporate the dollar cost per task to provide a more practical assessment for commercial applications.
When evaluating AI agents, the total cost of task completion is what matters. A model with a higher per-token cost can be more economical if it resolves a user's query in fewer turns than a cheaper, less capable model. This makes "number of turns" a primary efficiency metric.
Standardized benchmarks for AI models are largely irrelevant for business applications. Companies need to create their own evaluation systems tailored to their specific industry, workflows, and use cases to accurately assess which new model provides a tangible benefit and ROI.
Traditional AI benchmarks with percentage-based scores often saturate, losing their signal as models improve. Evals like VendingBench, which measure performance in dollars, have no upper ceiling. This provides a more durable and meaningful way to track AI progress and capabilities compared to finite scoring systems.
Traditional AI benchmarks are seen as increasingly incremental and less interesting. The new frontier for evaluating a model's true capability lies in applied, complex tasks that mimic real-world interaction, such as building in Minecraft (MC Bench) or managing a simulated business (VendingBench), which are more revealing of raw intelligence.
OpenAI CEO Sam Altman advocates for evaluating AI models on "per-task pricing"—the total cost to achieve a desired outcome. This shifts focus from cheap per-token costs to overall efficiency, where a smarter, more expensive model can be cheaper for completing the final task.
Focusing on token pricing is misleading. A more powerful model may be more expensive per token but significantly cheaper per task because its higher efficiency requires fewer prompts and iterations to achieve a final result. The correct way to measure cost-effectiveness is by the total cost to complete a job, not the price of the raw material.
A cheap model that fails often becomes expensive due to retries, fallbacks, and human review. The true measure of economic efficiency is the cost to reliably complete a task, not the raw inference cost, which can be a misleading metric at scale.
The rapid release of new AI models makes it crucial for companies to move beyond industry benchmarks. Developing internal evaluation systems ("evals") is necessary to test and determine which model performs best for unique, high-value business use cases, as model choice is becoming extremely important.