We scan new podcasts and send you the top 5 insights daily.
AI can generate rule-based "top-down" evaluations from a task description (e.g., word count). However, discovering nuanced "bottom-up" evaluations requires human intuition from reviewing many real-world data samples to find subtle, recurring failure modes.
Current AI excels at information gathering, similar to a junior analyst. However, it lacks the meta-level learning to develop true expertise from repeated tasks. This makes it a powerful tool for amplifying existing experts by handling tedious work, not replacing their decision-making capabilities.
Generic evaluation metrics like "helpfulness" or "conciseness" are vague and untrustworthy. A better approach is to first perform manual error analysis to find recurring problems (e.g., "tour scheduling failures"). Then, build specific, targeted evaluations (evals) that directly measure the frequency of these concrete issues, making metrics meaningful.
Despite AI's power, it cannot replace the human element of data analysis, which requires stakeholder management, domain knowledge, and critical thinking to validate results. An AI can produce errors, making human judgment more crucial than ever to avoid costly mistakes and provide true insights.
Don't ask an LLM to perform initial error analysis; it lacks the product context to spot subtle failures. Instead, have a human expert write detailed, freeform notes ("open codes"). Then, leverage an LLM's strength in synthesis to automatically categorize those hundreds of human-written notes into actionable failure themes ("axial codes").
AI agents can flawlessly execute predefined tasks (SOPs). However, they still require significant human management to ensure high-quality output, apply taste, and surface meaningful signals from the data they generate. This creates a new layer of human work, rather than a complete replacement.
Automated evaluation platforms are effective at spotting clear failures, like a tool-call error. However, they consistently miss subtle problems that require deep product judgment and domain expertise, such as a sales bot mishandling a customer's objection.
When implementing AI for business use, the knowledge needed to evaluate models resides in subjective human experience. A key bottleneck is converting this domain-specific expertise into a machine-readable format that can be used to reliably assess AI performance against real-world business needs.
A one-size-fits-all evaluation method is inefficient. Use simple code for deterministic checks like word count. Leverage an LLM-as-a-judge for subjective qualities like tone. Reserve costly human evaluation for ambiguous cases flagged by the LLM or for validating new features.
The most valuable evals aren't built with complex software but are often simple spreadsheets. Their power comes from deep subject matter expertise, which is necessary to create nuanced prompts and accurate scoring criteria that truly test a model's ability in a specific domain like clinical genomics or law.
AI tools can dramatically accelerate test execution but lack the contextual understanding to interpret results or assess business risk. An effective hybrid model has humans own the 'what' and 'why' (sense-making) while AI handles the 'how fast' (execution).