Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

While AI excels at protein modeling thanks to direct data, "virtual cell" models are underperforming simple baselines. The core issue is their reliance on single-cell RNA-seq data, which acts as a poor, compressed representation of the cell's true, complex state, unlike the direct data available for proteins.

Related Insights

AI isn't just for designing RNA sequences. Its real value is in creating predictive models of complex cellular functions. This allows scientists to determine the precise set of instructions (RNAs) needed to make a cell perform a complex series of tasks, like targeting a brain tumor.

The primary bottleneck for creating powerful foundation models in biology is the lack of clean, large-scale experimental data—orders of magnitude less than what's available for LLMs. This creates a major opportunity for "data foundries" that use robotic labs to generate high-quality biological data at scale.

To create a predictive "virtual cell," data collection must shift from passive observation to active intervention. The strategy is to massively scale perturbation experiments (like Perturb-seq) across countless contexts and measure multi-modal responses, teaching the model cause and effect.

Instead of pursuing a purely academic goal of simulating every biochemical process, Noetik's "virtual cell" models are practical tools. They focus on understanding cell biology through heuristics that are useful for making drugs, like predicting a cell's transcriptome or protein expression in a specific context.

Today's "virtual cell" models represent training data well but cannot predict outcomes for novel interventions. The next frontier is building models that generalize to serve as true predictive oracles for experiments that haven't yet been performed, a key focus for BioHub.

AI models trained on descriptive data (e.g., RNA-seq) can classify cell states but fail to predict how to transition a diseased cell to a healthy one. True progress requires generating massive "causal" datasets that show the effects of specific genetic perturbations.

The primary obstacle to creating sophisticated AI models of cells isn't the AI itself, but the data. Existing datasets often perturb only one cellular variable at a time, failing to capture the complex interactions that arise from simultaneous changes. New platforms are needed to generate this multi-dimensional data.

The progress of AI in predicting cancer treatment is stalled not by algorithms, but by the data used to train them. Relying solely on static genetic data is insufficient. The critical missing piece is functional, contextual data showing how patient cells actually respond to drugs.

A major misconception is that general-purpose Large Language Models (LLMs) can be readily applied to complex biological problems. Biological data, like RNA sequencing, constitutes a unique language that requires custom-built foundation models, not simply fine-tuning of existing LLMs.

Genomic data (DNA) provides a static blueprint of potential, not a view of the actual biological activity. True understanding requires measuring the dynamic interactions of molecules and cells within tissues "downstream." Current methods capture only fragmentary slices, missing the full picture.