We scan new podcasts and send you the top 5 insights daily.
Foundation models like OpenAI prioritize historical, large-scale literature over newer, smaller medical studies. In under-researched domains such as hormone therapy and female testosterone use, this training hierarchy leads to outdated or inaccurate guidance. Relying on raw consumer LLMs for specialized healthcare can actively harm patients by reinforcing outdated clinical biases rather than reflecting recent clinical findings.
An experiment using two leading AI models (Copilot and Gemini) to summarize 15 publications yielded contradictory and incomplete results. This demonstrates that relying on AI output without rigorous human verification can lead to dangerously misinformed conclusions in medical communications.
While powerful for analyzing existing medical data, AI struggles with true scientific discovery where the underlying biological principles are still unknown. Since AI learns from existing data, it cannot easily generate hypotheses that violate the very rules it was trained on, limiting its role in frontier science.
Despite the hype, Datycs' CEO finds that even fine-tuned healthcare LLMs struggle with the real-world complexity and messiness of clinical notes. This reality check highlights the ongoing need for specialized NLP and domain-specific tools to achieve accuracy in healthcare.
When a lab report screenshot included a dismissive note about "hemolysis," both human doctors and a vision-enabled AI made the same mistake of ignoring a critical data point. This highlights how AI can inherit human biases embedded in data presentation, underscoring the need to test models with varied information formats.
The danger of LLMs in research extends beyond simple hallucinations. Because they reference scientific literature—up to 50% of which may be irreproducible in life sciences—they can confidently present and build upon flawed or falsified data, creating a false sense of validity and amplifying the reproducibility crisis.
General-purpose LLMs generate responses based on the average of vast datasets. When used for leadership advice, they risk promoting a 'median' or average leadership style. This not only stifles authenticity but can also reinforce historical biases present in the training data.
The focus on preventing major, catastrophic AI errors overlooks the more pervasive risk of subtle misalignment. This includes models making decisions based on hospital profitability rather than patient well-being, systematically degrading care without a single, obvious failure. This subtle bias is harder to define and detect.
Large biopharma companies have failed when attempting to use generalist large language models (LLMs) for deal scouting. These models lack the specialized focus and curated data required for the industry, leading to inaccurate results, disappointment, and ultimately, abandoned internal AI projects.
AI models trained on engagement metrics like citations might prioritize popular or sensationalist articles. This risks creating a feedback loop where less-cited but more fundamental research is ignored, potentially stifling long-term scientific discovery by creating an AI-driven popularity bias.
While AI cybersecurity is a concern, many MedTech innovators overlook a more fundamental danger: the AI model itself being flawed. An AI making a wrong recommendation, like a therapy app encouraging suicide, can have dire consequences without any malicious external actor involved.