Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Believing agents are a key user, the company built a tool that runs a suite of tests where an AI agent attempts to complete tasks using their product. This provides immediate, objective feedback on API design, SDK usability, and documentation clarity, replacing slow user testing.

Related Insights

The traditional product feedback loop is being compressed by AI. Instead of waiting for human developers to test a beta, companies like Stripe now see AI agents deployed instantly. These agents provide immediate, detailed feedback through logs, allowing for an unprecedented pace of iteration and development.

Instead of manual user testing, prompt an AI agent to adopt specific user personas, like a hurried product manager or a spec-focused engineer. The AI will then use your application from that persona's perspective, providing targeted, research-style feedback on friction points and user experience.

Go beyond rapid prototyping. AI workflows can instantly create a functional prototype and simultaneously generate a usability test to capture customer feedback. This closes the feedback loop, allowing you to synthesize results and build a V2 in a single session.

Dreamer's AI "Sidekick" builds apps using the same command-line interface available to human developers. This forced the team to create excellent documentation and a clear API surface, which not only enables the agent but also significantly improves the developer experience for humans, creating a virtuous cycle.

For tools designed for AI interaction, the ease with which an agent can use the product (AX) is as critical as the user experience (UX) for humans. This can be improved by directly asking the agent for feedback on how to make the product more ergonomic for it.

Notion treats its entire evaluation process as a coding agent problem. The system is designed for an agent to download a dataset, run an eval, identify a failure, debug the issue, and implement a fix, all within an automated loop. This turns quality assurance into a meta-problem for agents to solve.

The prompts for your "LLM as a judge" evals function as a new form of PRD. They explicitly define the desired behavior, edge cases, and quality standards for your AI agent. Unlike static PRDs, these are living documents, derived from real user data and are constantly, automatically testing if the product meets its requirements.

Evals shift product development from defining the 'how' to defining the 'what'. By creating quantifiable tests and success criteria, evals act like a modern PRD. This allows an AI model to creatively figure out the implementation while the team focuses on defining the desired outcome through concrete examples.

Moving beyond analytics, the company is developing an AI agent that navigates an application like a real person. This "AI personality" can identify and report on areas of friction it encounters, providing a new, automated method for product testing and user experience validation before real users struggle.

Go beyond basic tests by instructing the AI to visually inspect its work from a customer's perspective. Have it click through flows, check for confusing elements or low-trust signals, and verify the user experience. This transforms the AI from a simple code generator into an active QA and product tester.

Together AI Uses an 'Agent Evals' Tool to Continuously Test Product Design and Docs | RiffOn