Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Traditional PRDs struggle to describe the vast capabilities of GenAI products. Evals serve as a more effective specification, communicating desired product behavior through a set of concrete examples and expected outcomes, rather than abstract descriptions.

Related Insights

While evals involve testing, their purpose isn't just to report bugs (information), like traditional QA. For an AI PM, evals are a core tool to actively shape and improve the product's behavior and performance (transformation) by iteratively refining prompts, models, and orchestration layers.

Building non-deterministic AI products fundamentally changes the PM role. Instead of creating detailed, rigid specifications, the PM's primary task becomes defining and codifying "what good looks like." This is done by repeatedly grading AI outputs to train evaluation systems and guide the model's behavior.

When building an AI feature, a PM's first job isn't writing requirements; it's defining success. This is done by creating an eval set that translates abstract user desires (e.g., "beautiful kitchens" on Pinterest) into concrete, scorable examples, forming the product's true specification.

Evals transform product specs from ambiguous documents into testable, measurable criteria. This gives product managers more leverage and provides clear targets for engineers, improving alignment and the quality of the final product.

At Anthropic, the primary artifact for product managers is no longer the PRD. Instead, they create "evals" (evaluation sets) from user feedback to define problems and measure model improvements. This makes user needs directly actionable for AI researchers, changing a core PM workflow.

The primary bottleneck in improving AI is no longer data or compute, but the creation of 'evals'—tests that measure a model's capabilities. These evals act as product requirement documents (PRDs) for researchers, defining what success looks like and guiding the training process.

Instead of traditional product requirements documents, AI PMs should define success through a set of specific evaluation metrics. Engineers then work to improve the system's performance against these evals in a "hill climbing" process, making the evals the functional specification for the product.

Floto.ai uses a PXD, a spec written for both human engineers and AI coding agents. It moves beyond UI requirements to define the conversational experience with principles, guardrails ('what not to do'), and examples of good/bad interactions, effectively 'tuning' the agent's behavior.

The prompts for your "LLM as a judge" evals function as a new form of PRD. They explicitly define the desired behavior, edge cases, and quality standards for your AI agent. Unlike static PRDs, these are living documents, derived from real user data and are constantly, automatically testing if the product meets its requirements.

Evals shift product development from defining the 'how' to defining the 'what'. By creating quantifiable tests and success criteria, evals act like a modern PRD. This allows an AI model to creatively figure out the implementation while the team focuses on defining the desired outcome through concrete examples.

Evals Replace PRD 'How' Sections by Defining AI Behavior Through Examples | RiffOn