We scan new podcasts and send you the top 5 insights daily.
Don't just generate data. Strategically choose omics types based on a clear trade-off: cheap but less actionable (genomics) vs. expensive but highly actionable (fluxomics) or even interventional (CRISPR screens). This provides a practical framework for R&D.
The company's breakthrough potential comes not from collecting raw DNA, but from linking it at an individual level to a rich set of "phenotype" data, including proteomics, metabolomics, and transcriptomics. This deep, multi-layered dataset from novel populations is what unlocks actionable insights for drug discovery.
Acknowledging the "garbage in, garbage out" principle, Haya heavily invests in generating high-quality, layered, and paired multi-omic data from the same biological material. This curated input is considered the most critical component for building effective AI models to unlock new biology.
To avoid wasting limited funds, startups should first validate their target product profile with regulators and investors. This 'end in mind' approach allows them to work backward, defining the exact data packages needed and prioritizing only the experiments that directly contribute to that goal.
To create a predictive "virtual cell," data collection must shift from passive observation to active intervention. The strategy is to massively scale perturbation experiments (like Perturb-seq) across countless contexts and measure multi-modal responses, teaching the model cause and effect.
While genomics predicts lifelong risk, Regeneron was surprised to discover that proteomics provides a more powerful, dynamic snapshot of health. In many cases, an individual's proteome was more effective at predicting disease outcomes in the next one to five years than their inherited genome, prompting massive investment in the technology.
In biotech, early data is often ambiguous. Instead of judging programs on potential, leaders must prioritize based on the time and capital required to reach a clear 'yes' or 'no' outcome. Indefinite 'gray zone' projects drain resources that could fund a winner.
To truly understand biological systems, data scale is less important than data quality. The most informative data comes from capturing the dynamic interactions of a system *while* it's being perturbed (e.g., by a drug), not from static snapshots of a system at rest.
Unlike general AI which leverages vast, existing datasets, Noetik believes progress in biology requires designing and generating specific, high-quality data with foresight into the models that will be trained. They compare this to the intentional, decades-long creation of the PDB dataset for protein folding.
While petabytes of observational DNA sequence data exist, it's insufficient for the next wave of AI. The key to creating powerful, functional models is generating causal data—from experiments that systematically test function—which is a current data bottleneck.
Regeneron views genomics as a "blueprint" for long-term risk. In contrast, proteomics acts as a real-time "sensor" of the body's current state. Their research showed proteomic data was surprisingly more predictive than genetics for the near-term onset of hundreds of diseases, including cancer and heart disease.