Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Most diseases are linked to variants in non-coding DNA, which makes up 98% of the genome and is notoriously hard to analyze. New long-context AI models excel at detecting these long-range interactions, significantly outperforming older methods.

Related Insights

Pre-trained genomic models like EVO showed potential but were unaligned. By applying alignment techniques like mid-training and post-training—similar to turning a base LLM into a useful chatbot—the Omni model became state-of-the-art across multiple biological tasks.

By analyzing a model predicting Alzheimer's, Goodfire discovered it relied on the length of cell-free DNA fragments—a previously overlooked signal. This demonstrates how interpretability can extract new, testable scientific hypotheses from high-performing "black box" models.

Unsupervised genomic models learn the statistical patterns of healthy DNA. To predict if a mutation causes disease, they compare the probability (likelihood score) of the original sequence versus the mutated one. A large drop in probability signals a 'surprising' and likely pathogenic variant.

The progress of AI in predicting cancer treatment is stalled not by algorithms, but by the data used to train them. Relying solely on static genetic data is insufficient. The critical missing piece is functional, contextual data showing how patient cells actually respond to drugs.

The development of pioneering genomic models like HyenaDNA wasn't driven by a biological problem. It started with AI researchers developing efficient long-context architectures and then asking, 'What's the longest sequence data out there to test this on?' The answer was DNA.

While petabytes of observational DNA sequence data exist, it's insufficient for the next wave of AI. The key to creating powerful, functional models is generating causal data—from experiments that systematically test function—which is a current data bottleneck.

Frontier AI models excel in medicine less because of their encyclopedic knowledge and more because of their ability to integrate huge amounts of context. They can synthesize a patient's entire medical history with the latest research—a task difficult for any single human. This highlights that the key to unlocking AI's value is feeding it comprehensive data, as context is the primary driver of superhuman performance.

Simple linear models fail to generalize to new cell types because gene functions are highly context-dependent. Some "housekeeping" genes have universal effects, but many others behave differently in various cellular environments. A sophisticated, nonlinear AI model is required to capture these context-specific interactions.

Myome and Natera are building foundational models for oncology that function like genomic language models. By training on vast cancer sequence and clinical data, these models learn the context of a patient's disease to predict the next mutation, similar to how transformers like GPT predict the next word in a sentence.

A major frustration in genetics is finding 'variants of unknown significance' (VUS)—genetic anomalies with no known effect. AI models promise to simulate the impact of these unique variants on cellular function, moving medicine from reactive diagnostics to truly personalized, predictive health.