Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

In developing the X-Cell model, Xaira found a clear hierarchy of impact. The quality, scale, and causal nature of the training data provided the most significant performance boost, followed by the choice of AI architecture (e.g., diffusion vs. autoregressive), and lastly, the integration of prior biological knowledge.

Related Insights

The bottleneck for AI in drug discovery is not the algorithm but the lack of high-quality, large-scale biological data. New platforms are needed to generate this necessary "substrate" for AI models to learn from, challenging the narrative that better models alone are the solution.

Standard AI models trained on public, observational biological data excel at descriptive tasks but underperform even linear models on causal predictions. To predict cellular responses to drug-like perturbations, models must be trained specifically on causal data generated from targeted experiments.

Beyond its core architecture, X-Cell integrates five types of biological priors, including text embeddings from scientific literature, protein interaction networks, and morphology information. This diverse context allows the model to make more accurate predictions and provides interpretability by showing which priors are most important for specific cell types.

Unlike text-based LLMs where simply increasing parameter count works, Verge Labs found the biggest AI performance gains in biology come from scaling data modalities—adding new types of data like proteomics and imaging. Fusing different data sources is more critical than just making the model bigger.

Xaira's core strategy involves creating massive, proprietary datasets that reveal causal biology. By systematically perturbing every gene in a cell to observe its effects, they generate unique training data for their models, quadrupling the world's supply of such information with a single publication.

AI models trained on descriptive data (e.g., RNA-seq) can classify cell states but fail to predict how to transition a diseased cell to a healthy one. True progress requires generating massive "causal" datasets that show the effects of specific genetic perturbations.

The primary obstacle to creating sophisticated AI models of cells isn't the AI itself, but the data. Existing datasets often perturb only one cellular variable at a time, failing to capture the complex interactions that arise from simultaneous changes. New platforms are needed to generate this multi-dimensional data.

The bottleneck for AI in drug development isn't the sophistication of the models but the absence of large-scale, high-quality biological data sets. Without comprehensive data on how drugs interact within complex human systems, even the best AI models cannot make accurate predictions.

Unlike general AI which leverages vast, existing datasets, Noetik believes progress in biology requires designing and generating specific, high-quality data with foresight into the models that will be trained. They compare this to the intentional, decades-long creation of the PDB dataset for protein folding.

While petabytes of observational DNA sequence data exist, it's insufficient for the next wave of AI. The key to creating powerful, functional models is generating causal data—from experiments that systematically test function—which is a current data bottleneck.