Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Eric Nguyen recalls that when he first pitched the idea for the Evo model, established scientists were highly skeptical. Their logic was that since humans haven't deciphered all the rules of the genome, it would be impossible for an AI to learn them well enough to generate functional DNA.

Related Insights

The next major AI breakthrough will come from applying generative models to complex systems beyond human language, such as biology. By treating biological processes as a unique "language," AI could discover novel therapeutics or research paths, leading to a "Move 37" moment in science.

While powerful for analyzing existing medical data, AI struggles with true scientific discovery where the underlying biological principles are still unknown. Since AI learns from existing data, it cannot easily generate hypotheses that violate the very rules it was trained on, limiting its role in frontier science.

AI cannot yet revolutionize drug discovery because its strength is synthesizing existing knowledge. The problem is that humans only understand about 20% of the human body's biology, meaning the foundational dataset is too incomplete for AI to reliably predict outcomes for the unknown 80%.

The development of pioneering genomic models like HyenaDNA wasn't driven by a biological problem. It started with AI researchers developing efficient long-context architectures and then asking, 'What's the longest sequence data out there to test this on?' The answer was DNA.

Despite AI's rapid progress, David Sinclair states that fully simulating a single biological cell from the atomic level is beyond near-future computing. The quantum effects and sheer number of molecular interactions present a challenge that will likely require quantum computers.

Unlike math or code with cheap, fast rewards, clinically valuable biology problems lack easily verifiable ground truths. This makes it difficult to create the rapid reinforcement learning loops that drive explosive AI progress in other fields.

Current AIs are trained on the established, consensus-driven scientific literature. The real breakthrough will occur when AI is trained on the 'trash can corpus'—all the ideas and papers that were rejected, laughed at, and dismissed by the orthodoxy. This is where undiscovered alpha lies.

John Jumper uses an analogy to explain the leap in complexity from prediction to design. Predicting a protein's structure is like recognizing a bicycle's parts. Designing a new, functional protein is like building a working bicycle—requiring every detail to be correct.

Unlike general AI which leverages vast, existing datasets, Noetik believes progress in biology requires designing and generating specific, high-quality data with foresight into the models that will be trained. They compare this to the intentional, decades-long creation of the PDB dataset for protein folding.

While petabytes of observational DNA sequence data exist, it's insufficient for the next wave of AI. The key to creating powerful, functional models is generating causal data—from experiments that systematically test function—which is a current data bottleneck.