We scan new podcasts and send you the top 5 insights daily.
The low-hanging fruit of applying AI to existing datasets is being picked. The next major leap forward will come not from slightly better models, but from creative strategies to generate entirely new datasets for unsolved problems like protein stability or in vivo effects.
The bottleneck for AI in drug discovery is not the algorithm but the lack of high-quality, large-scale biological data. New platforms are needed to generate this necessary "substrate" for AI models to learn from, challenging the narrative that better models alone are the solution.
The next leap in biotech moves beyond applying AI to existing data. CZI pioneers a model where 'frontier biology' and 'frontier AI' are developed in tandem. Experiments are now designed specifically to generate novel data that will ground and improve future AI models, creating a virtuous feedback loop.
The primary bottleneck for creating powerful foundation models in biology is the lack of clean, large-scale experimental data—orders of magnitude less than what's available for LLMs. This creates a major opportunity for "data foundries" that use robotic labs to generate high-quality biological data at scale.
To break the data bottleneck in AI protein engineering, companies now generate massive synthetic datasets. By creating novel "synthetic epitopes" and measuring their binding, they can produce thousands of validated positive and negative training examples in a single experiment, massively accelerating model development.
Unlike text-based LLMs where simply increasing parameter count works, Verge Labs found the biggest AI performance gains in biology come from scaling data modalities—adding new types of data like proteomics and imaging. Fusing different data sources is more critical than just making the model bigger.
The next inflection point will come from clever data generation strategies optimized for AI models, not human analysis. This "black box data" approach—like pooled screening with sequencing readouts—is vastly more scalable and creates a powerful, proprietary moat for companies.
Unlike language models trained on existing internet data, Biohub's biological models require data that doesn't exist yet. Their strategy pairs a frontier AI lab with a "frontier biology" effort to invent new imaging and measurement tools, creating proprietary data streams to fuel their models.
The primary obstacle to creating sophisticated AI models of cells isn't the AI itself, but the data. Existing datasets often perturb only one cellular variable at a time, failing to capture the complex interactions that arise from simultaneous changes. New platforms are needed to generate this multi-dimensional data.
Unlike general AI which leverages vast, existing datasets, Noetik believes progress in biology requires designing and generating specific, high-quality data with foresight into the models that will be trained. They compare this to the intentional, decades-long creation of the PDB dataset for protein folding.
While petabytes of observational DNA sequence data exist, it's insufficient for the next wave of AI. The key to creating powerful, functional models is generating causal data—from experiments that systematically test function—which is a current data bottleneck.