Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

The healthcare AI startup uses a unique architecture where 30 LLMs supervise one patient-facing LLM. This is combined with extensive "output testing" by thousands of clinicians to ensure the AI's responses are safe before deployment, a more rigorous method than just evaluating training data.

Related Insights

Rather than relying on a small group of experts, OpenAI has built a three-tiered system involving over 260 physicians. This includes high-level strategic advisors, a large cohort for data operations like red-teaming and comparison tasks (communicating via Slack), and a core group of close advisors who translate this collective expertise into concrete evals and training data for researchers.

Unlike general enterprise AI where a wrong answer is an inconvenience, errors in healthcare AI can be fatal. This high-stakes environment forces companies like Abridge to adopt extremely rigorous offline evaluation and phased, progressive rollouts, a far more cautious approach than typical "move fast" software development.

You can't just deploy a probabilistic model like an LLM in a high-stakes field like healthcare. The key is to build a deterministic infrastructure (e.g., a rules engine with clinical guidelines) that governs the AI's operation, ensuring it operates safely within predefined constraints.

To manage compliance risk in regulated industries, treat AI agents like new employees. Before deployment, the agent must pass the same knowledge assessment a human would take. This quantifies the risk, turning a 'black box' AI into an observable and testable system with a verifiable accuracy score.

OpenAI's health division serves a dual purpose: delivering societal benefits and providing a real-world, high-stakes environment for AI safety research. Problems like scalable oversight (supervising superhuman AI) move from theoretical exercises to practical necessities when models outperform physicians on narrow tasks, creating concrete feedback loops that accelerate safety progress.

Lindy dramatically increases agent reliability with a "validator" system. Before an action is taken, a second LLM call acts as a judge, checking the proposed action against an extensive prompt or checklist. Even a simple "Are you sure?" prompt provides a significant reliability bump.

To ensure reliability in healthcare, ZocDoc doesn't give LLMs free rein. It wraps them in a hybrid system where traditional, deterministic code orchestrates the AI's tasks, sets firm boundaries, and knows when to hand off to a human, preventing the 'praying for the best' approach common with direct LLM use.

Healthcare is a model for AI governance beyond its regulatory framework. The industry has a pre-existing infrastructure of trust, experience with diverse use cases, established practices for post-deployment monitoring, and a deep understanding of human-in-the-loop systems, all directly applicable to AI.

Instead of relying solely on 'black box' LLMs, a more robust approach is neurosymbolic computation. This method combines three estimators: a traditional symbolic/rule-based model (e.g., a medical checklist), a neural network prediction, and an LLM's assessment. By comparing these diverse outputs, experts can make more informed and reliable judgments.

For high-stakes decisions like utilization management, validate an AI model by having it run alongside the existing human process. The AI renders a decision in parallel with the medical director, allowing the organization to confirm alignment and build confidence before “shifting left” to autonomous workflows.