Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Criticisms of AI "hallucinations" often miss the point. The proper benchmark for AI performance is not flawlessness but the alternative: a human analyst who also makes mistakes, gets tired, or uses poor sources. AI's tireless nature and the ability to run cheap, parallel checks can ultimately lead to higher reliability.

Related Insights

Current AI excels at information gathering, similar to a junior analyst. However, it lacks the meta-level learning to develop true expertise from repeated tasks. This makes it a powerful tool for amplifying existing experts by handling tedious work, not replacing their decision-making capabilities.

When discussing AI risks like hallucinations, former Chief Justice McCormack argues the proper comparison isn't a perfect system, but the existing human one. Humans get tired, biased, and make mistakes. The question isn't whether AI is flawless, but whether it's an improvement over the error-prone reality.

Demis Hassabis likens current AI models to someone blurting out the first thought they have. To combat hallucinations, models must develop a capacity for 'thinking'—pausing to re-evaluate and check their intended output before delivering it. This reflective step is crucial for achieving true reasoning and reliability.

The benchmark for AI performance shouldn't be perfection, but the existing human alternative. In many contexts, like medical reporting or driving, imperfect AI can still be vastly superior to error-prone humans. The choice is often between a flawed AI and an even more flawed human system, or no system at all.

AI's occasional errors ('hallucinations') should be understood as a characteristic of a new, creative type of computer, not a simple flaw. Users must work with it as they would a talented but fallible human: leveraging its creativity while tolerating its occasional incorrectness and using its capacity for self-critique.

Don't blindly trust AI. The correct mental model is to view it as a super-smart intern fresh out of school. It has vast knowledge but no real-world experience, so its work requires constant verification, code reviews, and a human-in-the-loop process to catch errors.

Advanced AI tools like "deep research" models can produce vast amounts of information, like 30-page reports, in minutes. This creates a new productivity paradox: the AI's output capacity far exceeds a human's finite ability to verify sources, apply critical thought, and transform the raw output into authentic, usable insights.

The benchmark for AI reliability isn't 100% perfection. It's simply being better than the inconsistent, error-prone humans it augments. Since human error is the root cause of most critical failures (like cyber breaches), this is an achievable and highly valuable standard.

Dr. Wachter warns that public perception will unfairly judge AI errors against an impossible standard of perfection, not against the flawed human alternative. A single AI mistake will be magnified, overshadowing its superior overall safety record and risking a backlash that stalls progress in healthcare.

Both humans and AI make mistakes. Instead of claiming AI is perfect, a more effective argument in regulated fields is that AI makes fewer mistakes and helps humans catch their own errors more quickly. This shifts the focus from perfection to improved safety and efficiency.