AI models capable of designing biological sequences are advancing rapidly. The same models can be used for defense (e.g., detecting pathogens), but this defensive side is significantly behind, creating a dangerous imbalance that needs to be addressed.
The underlying AI architecture for creating novel biological sequences is also highly effective at identifying dangerous ones. This dual-use nature means the team building design capabilities is best suited to build defense tools, as the models are fundamentally the same.
Pre-trained genomic models like EVO showed potential but were unaligned. By applying alignment techniques like mid-training and post-training—similar to turning a base LLM into a useful chatbot—the Omni model became state-of-the-art across multiple biological tasks.
By providing a genomic model with a sequence of examples showing progressively higher fitness scores (e.g., better RNA aptamers), the model learns the optimization trajectory. It can then continue this "thought process" to generate novel, even higher-performing sequences.
Since DNA is the source code for RNA and proteins, a foundation model pre-trained on DNA can learn underlying biological principles that transfer across modalities. This allows a single model to tackle tasks that previously required specialized protein or RNA models.
Most diseases are linked to variants in non-coding DNA, which makes up 98% of the genome and is notoriously hard to analyze. New long-context AI models excel at detecting these long-range interactions, significantly outperforming older methods.
Unsupervised genomic models learn the statistical patterns of healthy DNA. To predict if a mutation causes disease, they compare the probability (likelihood score) of the original sequence versus the mutated one. A large drop in probability signals a 'surprising' and likely pathogenic variant.
Existing bio-defense systems work by matching DNA sequences to known pathogen databases. However, generative AI can create novel sequences with different 'spellings' but the same dangerous function. Effective future defense must evolve to predict a sequence's function, not just its identity.
The development of pioneering genomic models like HyenaDNA wasn't driven by a biological problem. It started with AI researchers developing efficient long-context architectures and then asking, 'What's the longest sequence data out there to test this on?' The answer was DNA.
Eric Nguyen recalls that when he first pitched the idea for the Evo model, established scientists were highly skeptical. Their logic was that since humans haven't deciphered all the rules of the genome, it would be impossible for an AI to learn them well enough to generate functional DNA.
Deep expertise can sometimes lead to pessimism about new approaches. The key to breakthrough innovation is to collaborate with domain experts who possess deep knowledge but also maintain an optimistic and imaginative outlook, willing to challenge the status quo.
Top AI researchers often feel they must choose between working on frontier AI (e.g., chatbots) or applying AI to science. Radical Numerics argues you can do both, as biology presents fundamental challenges that push the boundaries of AI itself, from architecture to alignment.
