AI can generate rule-based "top-down" evaluations from a task description (e.g., word count). However, discovering nuanced "bottom-up" evaluations requires human intuition from reviewing many real-world data samples to find subtle, recurring failure modes.
When an LLM is given a long list of evaluation criteria, it may ignore some or get "lazy." A more robust approach is to use sub-agents, assigning each one a single, specific criterion to evaluate, ensuring thorough and reliable assessment.
Instead of using generic tools like spreadsheets for error analysis, leverage an AI agent to build a custom HTML interface. The agent analyzes your data's structure and renders it with visual encodings that make it far easier for a human to review and spot issues.
A powerful workflow for error analysis is an interactive loop. A human provides open-ended feedback on data samples in a custom UI. In the background, an AI agent monitors these interactions, distills them into themes, and proposes structured rubric criteria, effectively scaling human taste.
Automated evaluation platforms are effective at spotting clear failures, like a tool-call error. However, they consistently miss subtle problems that require deep product judgment and domain expertise, such as a sales bot mishandling a customer's objection.
To create a self-improving system, establish a loop where after you manually refine an AI's output, you prompt it to reflect on the entire conversation. Ask it to suggest specific updates to its own underlying skill and evaluations to avoid the same manual corrections in the future.
Getting an LLM to write in a way that a specific user finds personally satisfying is described as the "final loss" problem. This reflects the immense difficulty of capturing subjective taste, nuance, and an authentic individual voice, which goes far beyond simple factual accuracy.
Rather than reviewing random data samples, use an AI agent to first cluster the entire dataset. The agent can then select a diverse set of examples from across these clusters, ensuring the human reviewer is exposed to a wide range of behaviors and potential failures early in the process.
The crucial first step in building evaluations is not to start writing them immediately. Instead, begin by manually reviewing data and creating a structured approach to formalize and externalize your subjective taste and judgment. This provides a solid foundation for all subsequent evals.
