We scan new podcasts and send you the top 5 insights daily.
Contrary to past best practices, providing explicit examples within tool descriptions or system prompts can now degrade performance in advanced models. These models are imaginative enough to understand intention without being constrained by specific examples, which can negatively bias their output.
While detailed prompts are useful, starting with simple, open-ended prompts can unlock more creative and strategic responses from AI models. Experimenting with different levels of prompt detail across various models often yields surprising and superior results.
With models like Gemini 3, the key skill is shifting from crafting hyper-specific, constrained prompts to making ambitious, multi-faceted requests. Users trained on older models tend to pare down their asks, but the latest AIs are 'pent up with creative capability' and yield better results from bigger challenges.
For subjective tasks, refining instructions has diminishing returns. The most effective way to improve AI performance is to provide it with a set of high-quality examples of the desired output. A library of five great examples is more powerful than a perfectly crafted prompt.
For advanced AI models, providing a high-level goal rather than a detailed, prescriptive list of instructions often produces better outcomes. Over-prompting can constrain the model's intelligence, while a simpler prompt allows it to leverage its own planning capabilities for a more effective execution.
Instead of using absolute negatives like "never do X," explain the underlying reason you want to avoid X. This gives the model flexibility. For example, rather than a hard character limit, explaining the goal is a single tweet but allowing a thread if necessary gives the AI more freedom to create a better output.
Early AI tools required detailed, structured prompts with roles, examples, and constraints. Today's advanced models infer context, audience, and tone from simple instructions, shifting the required skill from "coaching" the AI to integrating its output effectively.
OpenAI found that removing repeated instructions from old prompts improved scores by 10-15% while cutting token usage by 66%. The complex rule lists built for older models now confuse systems like GPT-5.6, leading to worse and more expensive answers.
Comparing AI models based on single, identical prompts is a flawed methodology. A true evaluation involves 'driving' the model through multiple iterations of feedback and correction. This reveals its ability to understand and adapt to your specific intent, which is a far more critical measure of its utility than a single probabilistic output.
Good Star Labs found GPT-5's performance in their Diplomacy game skyrocketed with optimized prompts, moving it from the bottom to the top. This shows a model's inherent capability can be masked or revealed by its prompt, making "best model" a context-dependent title rather than an absolute one.
As AI models become more capable, overly detailed system prompts with many examples and hard constraints can be counterproductive. They limit the model's creativity. The Claude Code team cut their system prompt by 80% because the smarter model needs more freedom to find optimal solutions.