The next inflection point will come from clever data generation strategies optimized for AI models, not human analysis. This "black box data" approach—like pooled screening with sequencing readouts—is vastly more scalable and creates a powerful, proprietary moat for companies.
Retro Biosciences engineered proteins that reprogrammed aged cells 50 times more efficiently than standard methods. They achieved this in months by training a protein language model with OpenAI, compressing a process that took academics a decade and showcasing a dramatic acceleration in engineering timelines.
The biotech industry is held back because it lacks standardized, API-driven lab automation services akin to AWS for tech. Without this "cloud lab" infrastructure, each company must build its own from scratch, hindering progress and preventing the lean, rapid development seen in software.
Beyond novel drug creation, AI's biggest contribution may be democratization. Tools like ChatGPT give wet lab scientists direct access to most bioinformatics workflows without needing specialized teams, dramatically increasing productivity and the pace of discovery across the entire field.
Despite major technological advancements over decades, including genomics, CRISPR, and machine learning, drug approval odds have not improved. They remain stuck at 8-10%, suggesting these tools have maintained the pipeline but haven't yet broken the fundamental discovery bottleneck.
While no AI-discovered drugs are approved yet, the guest predicts a high probability of one entering clinical trials within the next year. Full approval is then estimated to take five to ten years, marking a significant milestone for the AI drug discovery field.
Human minds struggle to grasp the vast complexity of biological systems. The guest argues that AI is the natural language for biology, just as mathematics is for physics, because AI models can capture the intricate, interconnected dynamics that are beyond human intuition.
While AI excels at protein modeling thanks to direct data, "virtual cell" models are underperforming simple baselines. The core issue is their reliance on single-cell RNA-seq data, which acts as a poor, compressed representation of the cell's true, complex state, unlike the direct data available for proteins.
The low-hanging fruit of applying AI to existing datasets is being picked. The next major leap forward will come not from slightly better models, but from creative strategies to generate entirely new datasets for unsolved problems like protein stability or in vivo effects.
