Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Because traditional wet-lab techniques test only ~50 sequences weekly, Early instituted massively parallel reporter assays testing 250,000 barcoded sequences in a single batch. By funneling high-quality outlier data into DNA-trained large language models paired with predictive oracle filters, each laboratory-to-model cycle compounds predictive accuracy. This flywheel effectively treats DNA as a programmable language, mirroring how AlphaFold unlocked programmable proteins.

Related Insights

Most hard genetic cases stall at 'variants of uncertain significance' (VUS). Gamo Labs aims to solve this by creating a feedback loop: AI agents identify a VUS, robotic biology labs conduct experiments to determine its function, and the biological data is used to train and improve the AI models.

By labeling each cell with a unique DNA barcode, all clones can be grown together in a single, manufacturing-relevant bioreactor. This shifts the core challenge from laborious individual cell measurements to a high-throughput sequencing and data analysis task, dramatically increasing efficiency and data richness.

Pre-trained genomic models like EVO showed potential but were unaligned. By applying alignment techniques like mid-training and post-training—similar to turning a base LLM into a useful chatbot—the Omni model became state-of-the-art across multiple biological tasks.

The primary bottleneck for creating powerful foundation models in biology is the lack of clean, large-scale experimental data—orders of magnitude less than what's available for LLMs. This creates a major opportunity for "data foundries" that use robotic labs to generate high-quality biological data at scale.

Earli combines wet lab experiments with AI in a continuous feedback loop. They test massive libraries of synthetic DNA promoter sequences, feed the performance results into a Large Language Model (LLM), which then designs new, potentially more effective sequences. This iterative process rapidly optimizes their cancer-specific genetic switches.

High-throughput biology uses techniques like PerturbSeq to run thousands of genetic perturbation experiments simultaneously in a single "pool" of cells. This method is highly scalable and, crucially, avoids the batch effects that plague traditional experiments, creating clean, uniform data essential for training large-scale AI models.

Early efforts like the Human Cell Atlas were criticized as mere data collection ("stamp collecting"). However, the rise of LLMs provided the key to unlock this data's value, transforming vast, unstructured biological datasets into systems that generate scientific insights and move biology from discovery to engineering.

Since DNA is the source code for RNA and proteins, a foundation model pre-trained on DNA can learn underlying biological principles that transfer across modalities. This allows a single model to tackle tasks that previously required specialized protein or RNA models.

Building biologically relevant AI is not a one-off process. It demands a continuous "lab in the loop" system where wet lab experiments generate proprietary data to train models, whose outputs are then physically tested in the lab. This iterative feedback cycle constantly refines the model's predictive accuracy.

Myome and Natera are building foundational models for oncology that function like genomic language models. By training on vast cancer sequence and clinical data, these models learn the context of a patient's disease to predict the next mutation, similar to how transformers like GPT predict the next word in a sentence.