Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

To innovate quickly and safely with AI, Kavak treats evaluations as a first-class citizen. They invest equal engineering time and money into building robust evals as they do into the agents themselves. These "brakes" enable them to accelerate development while ensuring quality and business alignment, focusing on conversion, not vanity metrics.

Related Insights

While evals involve testing, their purpose isn't just to report bugs (information), like traditional QA. For an AI PM, evals are a core tool to actively shape and improve the product's behavior and performance (transformation) by iteratively refining prompts, models, and orchestration layers.

Before building an AI agent, product managers must first create an evaluation set and scorecard. This 'eval-driven development' approach is critical for measuring whether training is improving the model and aligning its progress with the product vision. Without it, you cannot objectively demonstrate progress.

AI models and frameworks change constantly. A deep understanding of user needs, encoded into a robust evaluation suite, is a lasting asset. This allows you to continuously iterate and improve quality, regardless of which new model or agent framework becomes popular.

Building reliable AI agents requires a developer mindset shift. The most critical task is not writing the agent's code but creating robust evaluations ('evals') that define and verify the desired business outcome. This makes a test-driven development approach non-negotiable for enterprise AI.

Companies like Revolut initially struggle with AI adoption, not due to technology, but because they must first build a "cold start" foundation of evaluation frameworks, metrics, and CI/CD for models. Once this is in place, their AI consumption grows exponentially, matching AI-native firms.

AI validation tools should be viewed as friction-reducers that accelerate learning cycles. They generate options, prototypes, and market signals faster than humans can. The goal is not to replace human judgment or predict success, but to empower teams to make better-informed decisions earlier.

The primary bottleneck in improving AI is no longer data or compute, but the creation of 'evals'—tests that measure a model's capabilities. These evals act as product requirement documents (PRDs) for researchers, defining what success looks like and guiding the training process.

Building a functional AI agent is just the starting point. The real work lies in developing a set of evaluations ("evals") to test if the agent consistently behaves as expected. Without quantifying failures and successes against a standard, you're just guessing, not iteratively improving the agent's performance.

Formal evaluations ("evals") are not effective for zero-to-one innovation. Anthropic's Thariq Shihipar advises early-stage startups to avoid building complex eval systems, which slow them down, and instead iterate fast to build intuition about what works. Evals are for scaling, not discovery.

Evals shift product development from defining the 'how' to defining the 'what'. By creating quantifiable tests and success criteria, evals act like a modern PRD. This allows an AI model to creatively figure out the implementation while the team focuses on defining the desired outcome through concrete examples.