Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Formal evaluations ("evals") are not effective for zero-to-one innovation. Anthropic's Thariq Shihipar advises early-stage startups to avoid building complex eval systems, which slow them down, and instead iterate fast to build intuition about what works. Evals are for scaling, not discovery.

Related Insights

Don't treat evals as a mere checklist. Instead, use them as a creative tool to discover opportunities. A well-designed eval can reveal that a product is underperforming for a specific user segment, pointing directly to areas for high-impact improvement that a simple "vibe check" would miss.

For early-stage AI companies, performance should be measured by the speed of iteration, shipping, and learning, not just traditional metrics like revenue. In a rapidly evolving landscape, the ability to quickly get signals from the market and adapt is the primary indicator of future success.

With AI, teams can create crude prototypes immediately after a customer call. This "build to learn" phase cheaply validates ideas. Only after confirming market need should teams shift to "build to earn," investing in scalable development. This strategy mitigates the risk of building unwanted products at high speed.

Unlike traditional software development, where consistency is paramount, AI development requires testing many ideas quickly. Anthropic intentionally launches overlapping features to see which form factor users prefer, accepting the cost of a less consistent UX in exchange for speed and market feedback.

A "vibe check" is simply using your brain as a scoring function to intuit if an AI output is good. This aligns with the "do things that don't scale" startup principle and is a necessary first step before building more robust, scalable evaluation systems.

Unlike traditional software development that starts with unit tests for quality assurance, AI product development often begins with 'vibe testing.' Developers test a broad hypothesis to see if the model's output *feels* right, prioritizing creative exploration over rigid, predefined test cases at the outset.

Building custom AI model evaluations is a late-stage concern for creating defensibility. For pre-seed startups, it's a distraction. The only goal is achieving product-market fit using the cheapest, most accessible models, whether that's OpenAI, Anthropic, or open-source alternatives.

To innovate at the speed of AI, adopt the mindset that anything you build today could be made obsolete by next week's model release. This forces you to hold ideas loosely, constantly update your beliefs, and prioritize learning and exploration over perfection.

Founders with low trial volume often mistakenly try to A/B test small changes. With insufficient data, such tests are meaningless. Instead, they should focus on making big, obvious improvements based on gut feel and qualitative feedback. At this stage, the goal isn't optimization; it's finding significant wins that don't require statistical validation.

Prioritize qualitative 'vibe testing' over quantitative evals in early agent development. The most crucial first step is getting the agent in front of users to see if it 'feels' right and is useful before investing in formal, scalable quality checks.