We scan new podcasts and send you the top 5 insights daily.
Current virtual cell models saturate at only 2% of input data, not because of a lack of data, but because the data quality is poor. Cells grown in a petri dish are in an unnatural state, solely focused on proliferation, which masks the true effects of genetic perturbations and makes most of the data uninformative.
While AI excels at protein modeling thanks to direct data, "virtual cell" models are underperforming simple baselines. The core issue is their reliance on single-cell RNA-seq data, which acts as a poor, compressed representation of the cell's true, complex state, unlike the direct data available for proteins.
The primary bottleneck for creating powerful foundation models in biology is the lack of clean, large-scale experimental data—orders of magnitude less than what's available for LLMs. This creates a major opportunity for "data foundries" that use robotic labs to generate high-quality biological data at scale.
To create a predictive "virtual cell," data collection must shift from passive observation to active intervention. The strategy is to massively scale perturbation experiments (like Perturb-seq) across countless contexts and measure multi-modal responses, teaching the model cause and effect.
Today's "virtual cell" models represent training data well but cannot predict outcomes for novel interventions. The next frontier is building models that generalize to serve as true predictive oracles for experiments that haven't yet been performed, a key focus for BioHub.
AI models trained on descriptive data (e.g., RNA-seq) can classify cell states but fail to predict how to transition a diseased cell to a healthy one. True progress requires generating massive "causal" datasets that show the effects of specific genetic perturbations.
The primary obstacle to creating sophisticated AI models of cells isn't the AI itself, but the data. Existing datasets often perturb only one cellular variable at a time, failing to capture the complex interactions that arise from simultaneous changes. New platforms are needed to generate this multi-dimensional data.
The progress of AI in predicting cancer treatment is stalled not by algorithms, but by the data used to train them. Relying solely on static genetic data is insufficient. The critical missing piece is functional, contextual data showing how patient cells actually respond to drugs.
To truly understand biological systems, data scale is less important than data quality. The most informative data comes from capturing the dynamic interactions of a system *while* it's being perturbed (e.g., by a drug), not from static snapshots of a system at rest.
Guardant's co-CEO is skeptical of many AI-in-biology efforts because the underlying public data is often 'very under sampled.' These low-resolution datasets miss the rare signals that differentiate cells, causing AI models to learn from noise rather than true biological drivers of disease.
To build an AI that can reason about biology, researchers must first create a massive "dictionary" of normal cellular behavior. This involves capturing ~50 petabytes of unperturbed dynamics in organisms like zebra fish. Only after establishing this baseline of "native dynamics" can the model effectively interpret data from perturbed or diseased states.