We scan new podcasts and send you the top 5 insights daily.
Beyond its core architecture, X-Cell integrates five types of biological priors, including text embeddings from scientific literature, protein interaction networks, and morphology information. This diverse context allows the model to make more accurate predictions and provides interpretability by showing which priors are most important for specific cell types.
In developing the X-Cell model, Xaira found a clear hierarchy of impact. The quality, scale, and causal nature of the training data provided the most significant performance boost, followed by the choice of AI architecture (e.g., diffusion vs. autoregressive), and lastly, the integration of prior biological knowledge.
AI isn't just for designing RNA sequences. Its real value is in creating predictive models of complex cellular functions. This allows scientists to determine the precise set of instructions (RNAs) needed to make a cell perform a complex series of tasks, like targeting a brain tumor.
Numenos AI found that unifying biological data without traditional borders, such as incorporating mouse data or cancer data for dermatological diseases, surprisingly increases the predictive accuracy of their models. This challenges the siloed approach to traditional research.
Unlike text-based LLMs where simply increasing parameter count works, Verge Labs found the biggest AI performance gains in biology come from scaling data modalities—adding new types of data like proteomics and imaging. Fusing different data sources is more critical than just making the model bigger.
The next frontier in preclinical research involves feeding multi-omics and spatial data from complex 3D cell models into AI algorithms. This synergy will enable a crucial shift from merely observing biological phenomena to accurately predicting therapeutic outcomes and patient responses.
Xaira's strategy combines three distinct AI platforms: one for protein design to create novel therapeutics, a "virtual cell" model to predict biological effects, and a patient representation model to predict clinical outcomes. This integrated approach aims to de-risk and accelerate the entire drug discovery pipeline.
Achieving explainability in AI for drug development isn't about post-hoc analysis. It requires building models from the ground up using inherently interpretable data like RNA sequencing and mutational profiles. When the inputs are explainable, the model's outputs become explainable by design.
Biohub applies mechanistic interpretability to its protein language models. By analyzing the model's internal representations—learned from both known and unknown biology—researchers can uncover emergent biological principles. This turns the model from a black box predictor into an engine for scientific discovery itself.
Simple linear models fail to generalize to new cell types because gene functions are highly context-dependent. Some "housekeeping" genes have universal effects, but many others behave differently in various cellular environments. A sophisticated, nonlinear AI model is required to capture these context-specific interactions.
The key utility of a "virtual cell" model isn't just predicting outcomes within its training data. Its power is the ability to generalize and make accurate causal predictions in entirely new contexts, such as different cell types or primary cells from donors, where large-scale experiments are difficult or impossible.