We scan new podcasts and send you the top 5 insights daily.
The development of pioneering genomic models like HyenaDNA wasn't driven by a biological problem. It started with AI researchers developing efficient long-context architectures and then asking, 'What's the longest sequence data out there to test this on?' The answer was DNA.
Top AI researchers often feel they must choose between working on frontier AI (e.g., chatbots) or applying AI to science. Radical Numerics argues you can do both, as biology presents fundamental challenges that push the boundaries of AI itself, from architecture to alignment.
The next major AI breakthrough will come from applying generative models to complex systems beyond human language, such as biology. By treating biological processes as a unique "language," AI could discover novel therapeutics or research paths, leading to a "Move 37" moment in science.
Pre-trained genomic models like EVO showed potential but were unaligned. By applying alignment techniques like mid-training and post-training—similar to turning a base LLM into a useful chatbot—the Omni model became state-of-the-art across multiple biological tasks.
The next leap in biotech moves beyond applying AI to existing data. CZI pioneers a model where 'frontier biology' and 'frontier AI' are developed in tandem. Experiments are now designed specifically to generate novel data that will ground and improve future AI models, creating a virtuous feedback loop.
Jensen Huang forecasts that the next major AI breakthrough will be in digital biology. He believes advances in multimodality, long context models, and synthetic data will converge to create a "ChatGPT moment," enabling the generation of novel proteins and chemicals.
By providing a genomic model with a sequence of examples showing progressively higher fitness scores (e.g., better RNA aptamers), the model learns the optimization trajectory. It can then continue this "thought process" to generate novel, even higher-performing sequences.
Eric Nguyen recalls that when he first pitched the idea for the Evo model, established scientists were highly skeptical. Their logic was that since humans haven't deciphered all the rules of the genome, it would be impossible for an AI to learn them well enough to generate functional DNA.
While petabytes of observational DNA sequence data exist, it's insufficient for the next wave of AI. The key to creating powerful, functional models is generating causal data—from experiments that systematically test function—which is a current data bottleneck.
Most diseases are linked to variants in non-coding DNA, which makes up 98% of the genome and is notoriously hard to analyze. New long-context AI models excel at detecting these long-range interactions, significantly outperforming older methods.
Myome and Natera are building foundational models for oncology that function like genomic language models. By training on vast cancer sequence and clinical data, these models learn the context of a patient's disease to predict the next mutation, similar to how transformers like GPT predict the next word in a sentence.