Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Unsupervised genomic models learn the statistical patterns of healthy DNA. To predict if a mutation causes disease, they compare the probability (likelihood score) of the original sequence versus the mutated one. A large drop in probability signals a 'surprising' and likely pathogenic variant.

Related Insights

While sequencing costs are plummeting, the true value lies in interpreting the data. HLI's competitive advantage is its AI model, trained on a proprietary, decade-long dataset linking genomics with deep phenotyping for over 10,000 clients.

Wei-Wu He: Craig Venter’s Legacy and the Future of Human Longevity thumbnail

Wei-Wu He: Craig Venter’s Legacy and the Future of Human Longevity

Behind the Breakthroughs·11 days ago

Training a language model to predict the next amino acid in a sequence forces it to learn the protein's 3D structure. To make accurate predictions, the model must understand an amino acid's physical microenvironment, effectively deriving 3D spatial relationships from 1D sequence data alone. This demonstrates emergent capabilities of LLMs in biology.

Pre-trained genomic models like EVO showed potential but were unaligned. By applying alignment techniques like mid-training and post-training—similar to turning a base LLM into a useful chatbot—the Omni model became state-of-the-art across multiple biological tasks.

By analyzing a model predicting Alzheimer's, Goodfire discovered it relied on the length of cell-free DNA fragments—a previously overlooked signal. This demonstrates how interpretability can extract new, testable scientific hypotheses from high-performing "black box" models.

Existing bio-defense systems work by matching DNA sequences to known pathogen databases. However, generative AI can create novel sequences with different 'spellings' but the same dangerous function. Effective future defense must evolve to predict a sequence's function, not just its identity.

A major misconception is that general-purpose Large Language Models (LLMs) can be readily applied to complex biological problems. Biological data, like RNA sequencing, constitutes a unique language that requires custom-built foundation models, not simply fine-tuning of existing LLMs.

Achieving explainability in AI for drug development isn't about post-hoc analysis. It requires building models from the ground up using inherently interpretable data like RNA sequencing and mutational profiles. When the inputs are explainable, the model's outputs become explainable by design.

Most diseases are linked to variants in non-coding DNA, which makes up 98% of the genome and is notoriously hard to analyze. New long-context AI models excel at detecting these long-range interactions, significantly outperforming older methods.

Myome and Natera are building foundational models for oncology that function like genomic language models. By training on vast cancer sequence and clinical data, these models learn the context of a patient's disease to predict the next mutation, similar to how transformers like GPT predict the next word in a sentence.

A major frustration in genetics is finding 'variants of unknown significance' (VUS)—genetic anomalies with no known effect. AI models promise to simulate the impact of these unique variants on cellular function, moving medicine from reactive diagnostics to truly personalized, predictive health.