/
© 2026 RiffOn. All rights reserved.

Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

  1. Behind the Craft
  2. How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel
How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel

Behind the Craft · Aug 23, 2026

Elevate your AI evals. Shreya & Hamel share a 5-step process to automate error analysis, combining human taste with AI-powered tools.

AI Excels at Top-Down Evals, But Humans Must Drive Bottom-Up Evals From Data

AI can generate rule-based "top-down" evaluations from a task description (e.g., word count). However, discovering nuanced "bottom-up" evaluations requires human intuition from reviewing many real-world data samples to find subtle, recurring failure modes.

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel thumbnail

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel

Behind the Craft·a month ago

Use Focused Sub-Agents to Evaluate Distinct Criteria and Prevent LLM Laziness

When an LLM is given a long list of evaluation criteria, it may ignore some or get "lazy." A more robust approach is to use sub-agents, assigning each one a single, specific criterion to evaluate, ensuring thorough and reliable assessment.

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel thumbnail

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel

Behind the Craft·a month ago

Use AI to Build Custom Data Review Interfaces That Surpass Spreadsheets

Instead of using generic tools like spreadsheets for error analysis, leverage an AI agent to build a custom HTML interface. The agent analyzes your data's structure and renders it with visual encodings that make it far easier for a human to review and spot issues.

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel thumbnail

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel

Behind the Craft·a month ago

Create an Interactive Loop Where AI Synthesizes Human Feedback into a Rubric in Real-Time

A powerful workflow for error analysis is an interactive loop. A human provides open-ended feedback on data samples in a custom UI. In the background, an AI agent monitors these interactions, distills them into themes, and proposes structured rubric criteria, effectively scaling human taste.

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel thumbnail

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel

Behind the Craft·a month ago

Automated Eval Tools Catch Obvious Errors But Miss Failures Requiring Product Taste

Automated evaluation platforms are effective at spotting clear failures, like a tool-call error. However, they consistently miss subtle problems that require deep product judgment and domain expertise, such as a sales bot mishandling a customer's objection.

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel thumbnail

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel

Behind the Craft·a month ago

Prompt Your AI to Suggest Updates For Its Own Skill and Evals After Each Use

To create a self-improving system, establish a loop where after you manually refine an AI's output, you prompt it to reflect on the entire conversation. Ask it to suggest specific updates to its own underlying skill and evaluations to avoid the same manual corrections in the future.

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel thumbnail

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel

Behind the Craft·a month ago

Writing in a User's Specific Style is the 'Final Loss' and Toughest Challenge for LLMs

Getting an LLM to write in a way that a specific user finds personally satisfying is described as the "final loss" problem. This reflects the immense difficulty of capturing subjective taste, nuance, and an authentic individual voice, which goes far beyond simple factual accuracy.

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel thumbnail

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel

Behind the Craft·a month ago

Use AI to Cluster Data and Select Diverse Samples for More Efficient Human Review

Rather than reviewing random data samples, use an AI agent to first cluster the entire dataset. The agent can then select a diverse set of examples from across these clusters, ensuring the human reviewer is exposed to a wide range of behaviors and potential failures early in the process.

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel thumbnail

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel

Behind the Craft·a month ago

Externalize Your Judgment and Product Taste Before Writing Any AI Evaluations

The crucial first step in building evaluations is not to start writing them immediately. Instead, begin by manually reviewing data and creating a structured approach to formalize and externalize your subjective taste and judgment. This provides a solid foundation for all subsequent evals.

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel thumbnail

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel

Behind the Craft·a month ago