Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

To improve generative UI, designers create a "golden set" of 10-20 ideal prompts and their desired outcomes. They repeatedly run the model against this set, identify failures, tweak the guiding "skill," and iterate, creating a flywheel for quality improvement.

Related Insights

Developing a high-quality AI skill, like an "Ad Optimizer," is not as simple as writing a single prompt. It requires a laborious, iterative cycle of instructing, testing, analyzing poor outputs, and refining the instructions—much like training a human employee. This effort will become a key differentiator.

For subjective tasks, refining instructions has diminishing returns. The most effective way to improve AI performance is to provide it with a set of high-quality examples of the desired output. A library of five great examples is more powerful than a perfectly crafted prompt.

Users mistakenly evaluate AI tools based on the quality of the first output. However, since 90% of the work is iterative, the superior tool is the one that handles a high volume of refinement prompts most effectively, not the one with the best initial result.

Instead of manually refining a complex prompt, create a process where an AI agent evaluates its own output. By providing a framework for self-critique, including quantitative scores and qualitative reasoning, the AI can iteratively enhance its own system instructions and achieve a much stronger result.

To manage non-deterministic AI products, Shopify created an internal tool where PMs grade AI-generated outputs. This creates a "ground truth" dataset of what "good" looks like, which is then used to fine-tune a separate LLM that acts as an automated quality judge for new features and updates.

Expecting employees to author perfect, complex prompts from scratch leads to paralysis. A better method is letting them complete a task via iteration with the AI, then having the system automatically capture those adjustments as a reusable workflow or 'skill.'

For generative UI, designers move from defining exact layouts to creating a "skill" that guides the AI. This skill defines a design language, references tokens, and suggests best practices, steering the model's output without over-constraining it.

Instead of perfecting a single prompt, treat AI interaction as a rapid, iterative cycle. View the first output as a draft. Like managing an employee, provide feedback and refine the result over several short cycles to achieve a superior outcome, which is more effective than front-loading all effort.

Many people struggle to define what 'good' looks like. Building an evaluation (eval) for an AI system requires you to codify your quality standards, forcing a level of clarity and commitment that improves your own process and the AI's output.

The latest AI models appear more creative in design tasks not by learning new skills, but by being explicitly trained to avoid generic outputs that signal AI generation, like "bento box" layouts. This strategy of identifying and eliminating "bad AI smell" is a novel approach to improving the perceived quality and sophistication of generative models.