Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

The crucial first step in building evaluations is not to start writing them immediately. Instead, begin by manually reviewing data and creating a structured approach to formalize and externalize your subjective taste and judgment. This provides a solid foundation for all subsequent evals.

Related Insights

A practical first step with AI isn't asking for answers, but improving your process. Write down how you typically make a specific decision—what data you look for, what counter-arguments—then prompt an LLM to identify blind spots and suggest improvements to your framework.

Standard benchmarks are insufficient. A more effective evaluation method is a hybrid approach, weighting a human's qualitative 'taste test' (e.g., 70%) more heavily than an LLM judge's automated score (e.g., 30%). This prioritizes subjective qualities like design, usability, and writing style.

'Taste' is a collection of specific preferences, not an abstract feeling. Document what makes an output 'good' by creating universal rules (e.g., 'write at a ninth-grade level,' 'avoid cheesy quotes,' 'no em dashes'). Feeding these documented rules to an AI transforms your subjective taste into repeatable instructions for consistent results.

Don't start building evaluations from a blank slate. Use an AI agent to analyze your production traces and automatically generate a baseline 'vibe eval.' This initial evaluation won't be perfect, but it provides a starting point for refinement and accelerates the improvement loop.

The common mistake in building AI evals is jumping straight to writing automated tests. The correct first step is a manual process called "error analysis" or "open coding," where a product expert reviews real user interaction logs to understand what's actually going wrong. This grounds your entire evaluation process in reality.

Do not blindly trust an LLM's evaluation scores. The biggest mistake is showing stakeholders metrics that don't match their perception of product quality. To build trust, first hand-label a sample of data with binary outcomes (good/bad), then compare the LLM judge's scores against these human labels to ensure agreement before deploying the eval.

Prioritize qualitative 'vibe testing' over quantitative evals in early agent development. The most crucial first step is getting the agent in front of users to see if it 'feels' right and is useful before investing in formal, scalable quality checks.

Many people struggle to define what 'good' looks like. Building an evaluation (eval) for an AI system requires you to codify your quality standards, forcing a level of clarity and commitment that improves your own process and the AI's output.

Simply asking an LLM to "judge" an output yields generic results. A useful LLM judge requires manually injecting your own taste by creating detailed rubrics with extensive examples of "good" and "bad" at a granular level, essentially brute-forcing your preferences into the model.

This framework demystifies building an eval. Define your input data (e.g., user queries), specify the task your AI performs (from an LLM call to a complex agent), and create scoring functions that normalize outputs to a 0-1 range for consistent comparison.