We scan new podcasts and send you the top 5 insights daily.
Unlike fields with vast training data like image recognition, effective drug discovery has too few successful examples for AI to learn from alone. To be useful, AI models must be explicitly taught the foundational principles and complex rules of medicinal chemistry that human experts use.
AI modeling transforms drug development from a numbers game of screening millions of compounds to an engineering discipline. Researchers can model molecular systems upfront, understand key parameters, and design solutions for a specific problem, turning a costly screening process into a rapid, targeted design cycle.
The bottleneck for AI in drug discovery is not the algorithm but the lack of high-quality, large-scale biological data. New platforms are needed to generate this necessary "substrate" for AI models to learn from, challenging the narrative that better models alone are the solution.
AI cannot yet revolutionize drug discovery because its strength is synthesizing existing knowledge. The problem is that humans only understand about 20% of the human body's biology, meaning the foundational dataset is too incomplete for AI to reliably predict outcomes for the unknown 80%.
Current AI for protein engineering relies on small public datasets like the PDB (~10,000 structures), causing models to "hallucinate" or default to known examples. This data bottleneck, orders of magnitude smaller than data used for LLMs, hinders the development of novel therapeutics.
Early AI drug discovery platforms built robust models but often failed to generate relevant outputs. Their lack of deep biological understanding led to flawed data collection and training sets, creating a "garbage in, garbage out" problem where models were disconnected from real-world biology.
Beyond accelerating timelines, AI's real value lies in its ability to design molecules for targets previously considered 'hard-to-drug.' These models operate on different principles than traditional lab methods and are indifferent to historical challenges, opening up entirely new therapeutic possibilities.
Achieving explainability in AI for drug development isn't about post-hoc analysis. It requires building models from the ground up using inherently interpretable data like RNA sequencing and mutational profiles. When the inputs are explainable, the model's outputs become explainable by design.
The current, tangible breakthrough for AI in drug discovery is not identifying completely novel biological targets. Instead, it's rapidly designing effective molecules for known targets that have historically been considered "undruggable," compressing years of screening work into a month.
The immediate goal for AI in drug design is finding initial "hits" for difficult targets. The true endgame, however, is to train models on manufacturability data—like solubility and stability—so they can generate molecules that are already optimized, drastically compressing the development timeline.
AI thrives on learning from the vast, structured data evolution provides for proteins. Molly Gibson explains that small molecules lack this clear "language" or evolutionary history. This fundamental data gap is a primary reason generative AI has been slower to transform small molecule drug discovery compared to biologics.