Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

The most valuable evals aren't built with complex software but are often simple spreadsheets. Their power comes from deep subject matter expertise, which is necessary to create nuanced prompts and accurate scoring criteria that truly test a model's ability in a specific domain like clinical genomics or law.

Related Insights

While AI can help draft "skills" (reusable prompts), research shows human-authored skills perform better. This highlights the value of domain expertise. Use AI as a starting point, but refine instructions with specific knowledge, templates, and context for optimal results.

Constructing a robust eval set involves a process akin to binary search. First, establish a performance floor with an easy, canonical task to ensure basic competency. Then, establish a ceiling with a frontier-level hard task. This maps the model's capabilities and helps fill in the gaps with medium-difficulty prompts.

Building an AI application is becoming trivial and fast ("under 10 minutes"). The true differentiator and the most difficult part is embedding deep domain knowledge into the prompts. The AI needs to be taught *what* to look for, which requires human expertise in that specific field.

To gauge an expert's (human or AI) true depth, go beyond recall-based questions. Pose a complex problem with multiple constraints, like a skeptical audience, high anxiety, and a tight deadline. A genuine expert will synthesize concepts and address all layers of the problem, whereas a novice will give generic advice.

AI evaluation shouldn't be confined to engineering silos. Subject matter experts (SMEs) and business users hold the critical domain knowledge to assess what's "good." Providing them with GUI-based tools, like an "eval studio," is crucial for continuous improvement and building trustworthy enterprise AI.

The primary bottleneck in improving AI is no longer data or compute, but the creation of 'evals'—tests that measure a model's capabilities. These evals act as product requirement documents (PRDs) for researchers, defining what success looks like and guiding the training process.

Product managers may lack the expertise to create comprehensive evals from scratch. A better approach is to generate initial outputs with a base model, have subject matter experts review them, and use their direct feedback to define what constitutes a failure. It's easier for experts to spot mistakes than to predict them.

As AI capabilities become commoditized, the key to superior output is the user's domain expertise. An expert with precise vocabulary can guide an AI to produce better results in one attempt than a novice can in many, because they can articulate the desired outcome more effectively.

When implementing AI for business use, the knowledge needed to evaluate models resides in subjective human experience. A key bottleneck is converting this domain-specific expertise into a machine-readable format that can be used to reliably assess AI performance against real-world business needs.

The most valuable AI systems are built by people with deep knowledge in a specific field (like pest control or law), not by engineers. This expertise is crucial for identifying the right problems and, more importantly, for creating effective evaluations to ensure the agent performs correctly.