/
© 2026 RiffOn. All rights reserved.

Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

  1. The Growth Podcast
  2. How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google
How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google

How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google

The Growth Podcast · Jul 28, 2026

Ex-Meta/Google PM Daniel McKinnon explains how frontier-lab quality 'evals' are replacing PRDs to define success for complex, agentic AI systems.

Building AI Evals is a Binary Search Between Easy and Hard Problems

Constructing a robust eval set involves a process akin to binary search. First, establish a performance floor with an easy, canonical task to ensure basic competency. Then, establish a ceiling with a frontier-level hard task. This maps the model's capabilities and helps fill in the gaps with medium-difficulty prompts.

How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google thumbnail

How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google

The Growth Podcast·2 months ago

Startups Have an Edge Building for AI by Avoiding Incumbent Reinvention

Large companies like Google and Meta must undergo a painful process of reinventing their "classic consumer software building factory" for the AI era. Startups have a key advantage: they can build AI-native processes and cultures from a blank slate, which is often easier than retrofitting a massive organization.

How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google thumbnail

How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google

The Growth Podcast·2 months ago

An Optimal AI Eval Set Should Yield a 25-50% Initial Success Rate

Evals are designed to guide improvement. A set that scores 100% is too easy and provides no optimization path. A 0% score is too hard and offers no signal. A 25-50% success rate creates a "Goldilocks" zone—a challenging but achievable target for engineering teams to iterate against.

How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google thumbnail

How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google

The Growth Podcast·2 months ago

For AI PMs, Evals Are the 'Day Zero' Task That Precedes a PRD

When building an AI feature, a PM's first job isn't writing requirements; it's defining success. This is done by creating an eval set that translates abstract user desires (e.g., "beautiful kitchens" on Pinterest) into concrete, scorable examples, forming the product's true specification.

How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google thumbnail

How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google

The Growth Podcast·2 months ago

Meta’s Founder-Led Culture Drives Aggression but Risks Ruthless Team Cuts

Unlike Google's consensus-driven approach, Meta's culture under Mark Zuckerberg is highly aggressive and founder-led. This empowers teams with resources and conviction but also creates a high-stakes environment where entire teams, like the Llama 4 team, can be abruptly cut for perceived failures.

How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google thumbnail

How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google

The Growth Podcast·2 months ago

Evals Replace PRD 'How' Sections by Defining AI Behavior Through Examples

Traditional PRDs struggle to describe the vast capabilities of GenAI products. Evals serve as a more effective specification, communicating desired product behavior through a set of concrete examples and expected outcomes, rather than abstract descriptions.

How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google thumbnail

How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google

The Growth Podcast·2 months ago

Frontier AI Evals Now Focus on Agentic Tasks, Not Simple Question-Answering

Early LLM benchmarks like MMLU tested question-answering, a now-saturated capability. Today, leading labs like Anthropic evaluate models on agentic tasks involving multi-step reasoning and tool use, reflecting the shift in AI applications from search replacement to automated workflows.

How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google thumbnail

How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google

The Growth Podcast·2 months ago

Deep Domain Expertise, Not Tooling, Is Key to Frontier-Quality AI Evals

The most valuable evals aren't built with complex software but are often simple spreadsheets. Their power comes from deep subject matter expertise, which is necessary to create nuanced prompts and accurate scoring criteria that truly test a model's ability in a specific domain like clinical genomics or law.

How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google thumbnail

How to Build Frontier-Lab Quality Evals with Daniel McKinnon, ex-PM at Meta, Google

The Growth Podcast·2 months ago