Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

Since DNA is the source code for RNA and proteins, a foundation model pre-trained on DNA can learn underlying biological principles that transfer across modalities. This allows a single model to tackle tasks that previously required specialized protein or RNA models.

Related Insights

Instead of building from scratch, ProPhet leverages existing transformer models to create unique mathematical 'languages' for proteins and molecules. Their core innovation is an additional model that translates between them, creating a unified space to predict interactions at scale.

Training a language model to predict the next amino acid in a sequence forces it to learn the protein's 3D structure. To make accurate predictions, the model must understand an amino acid's physical microenvironment, effectively deriving 3D spatial relationships from 1D sequence data alone. This demonstrates emergent capabilities of LLMs in biology.

Pre-trained genomic models like EVO showed potential but were unaligned. By applying alignment techniques like mid-training and post-training—similar to turning a base LLM into a useful chatbot—the Omni model became state-of-the-art across multiple biological tasks.

The core philosophy behind ESMFold is that massive datasets and large transformer models can learn fundamental biological principles without needing built-in domain knowledge, applying Rich Sutton's "The Bitter Lesson" directly to bioinformatics.

Unlike text-based LLMs where simply increasing parameter count works, Verge Labs found the biggest AI performance gains in biology come from scaling data modalities—adding new types of data like proteomics and imaging. Fusing different data sources is more critical than just making the model bigger.

Biohub's goal was to create a general world model that "understands proteins." An emergent property of this generalist model was state-of-the-art performance in the highly specialized task of designing single-chain antibodies, a critical function for therapeutics. This demonstrates the power of general models to solve niche problems without explicit training.

The success of protein language models can be explained by Zellig Harris's 1954 linguistic theory. Just as a word's meaning is defined by its contexts, an amino acid's biological role is determined by the sequences it can appear in. The model learns this deep statistical structure, effectively learning biology.

A major misconception is that general-purpose Large Language Models (LLMs) can be readily applied to complex biological problems. Biological data, like RNA sequencing, constitutes a unique language that requires custom-built foundation models, not simply fine-tuning of existing LLMs.

Trained only on sequence prediction, ESM-C independently developed a hierarchical feature space mirroring decades of human scientific discovery. Its learned representations range from basic biochemical properties to complex, abstract functional concepts, all without prior biological knowledge.

Generate Biomedicines' AI learns the fundamental rules of protein structure and function, much like a language's grammar. This allows it to design entirely new proteins by generating novel "sentences" (sequences) that are biologically coherent and functional, rather than just mimicking existing ones found in nature.