We scan new podcasts and send you the top 5 insights daily.
In clinical trials, Google's LLM achieved higher first-guess diagnostic accuracy (91% vs 77%) than human doctors. More surprisingly, it was rated as more empathetic (82% vs 71%) by physician graders, suggesting AI can excel in the "soft skills" of patient care.
Contrary to popular belief, research shows AI models perform better than the average human on emotional intelligence tests. This capability allows AI to craft empathetic and emotionally resonant communications, even for sensitive topics like delivering bad news, often better than human counterparts.
In a partnership with Kenya's Penda Health, OpenAI conducted the first randomized controlled trial of an LLM co-pilot for physicians. The study demonstrated a statistically significant improvement in diagnosis and treatment outcomes for patients whose doctors used the AI assistant. This provides crucial real-world evidence that AI can move beyond lab benchmarks to tangibly improve care.
The widespread use of AI for health queries is set to change doctor visits. Patients will increasingly arrive with AI-generated analyses of their lab results and symptoms, turning appointments into a three-way consultation between the patient, the doctor, and the AI's findings, potentially improving diagnostic efficiency.
In a sign of recursive capability improvement, OpenAI found that its model-based grader for the HealthBench evaluation benchmark was more accurate and consistent than the average human physician performing the same grading task. This demonstrates that models can not only perform a task but also evaluate that performance at a superhuman level, a key component of scalable oversight.
AI only imitates empathy, but it can be more effective than human-delivered empathy in high-stress roles. AI has infinite patience and isn't burdened by emotional fatigue that affects professionals like doctors or paramedics, leading to an experience where the recipient feels more cared for.
A Google study revealed that while an AI's treatment plans were rated 98% appropriate by the third visit, human doctors' appropriateness declined after the first. This indicates humans may be prone to confirmation bias or premature diagnostic closure, a flaw that learning models overcome.
While the caring economy is often cited as a future source of human jobs, AI's ability to be infinitely patient gives it an "unfair advantage" in roles like medicine and teaching. AI doctors already receive higher ratings for bedside manner, challenging the assumption that these roles are uniquely human.
An AI saying 'I'm sorry you're sick' is executing a learned linguistic pattern. It lacks the biological underpinnings of genuine empathy: memory of pain, mirror neurons, and an internal state of concern. Users should not confuse a correct response with a shared emotional experience.
In studies where clinical psychologists evaluate anonymized transcripts, AI-generated therapy responses are often rated higher than human ones. This suggests AI's significant potential in mental health, particularly for increasing access to care.
As AI doctors consistently outperform humans in accuracy, the legal and ethical standard of care will shift. A human doctor ignoring a correct AI diagnosis that leads to patient harm could become a clear case of malpractice, forcing universal adoption of AI as a diagnostic partner.