Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

To reach its goal of 20 million sequenced individuals, Regeneron plans to use tokenization to link de-identified EHR data from one source (like a hospital) with biospecimens from another (like LabCorp). This strategy moves beyond single-institution collaborations to massively scale its ability to create linked datasets.

Related Insights

Regeneron's RGC is exploring a new business model beyond its internal R&D function. It plans to partner with direct-to-consumer (DTC) platforms to bring its genomic insights on health and wellness directly to patients, signaling an evolution from a pure data engine to a broader life sciences intelligence player.

Regeneron identified the main constraint in drug discovery as a lack of validated targets, not a shortage of advanced therapeutic tools. Their genetics engine was created to explore the 90% of the human genome that was untargeted by existing or experimental medicines, aiming to solve this core problem.

Regeneron's Genetics Center is a key competitive advantage, functioning as a discovery engine for new drug targets. By sequencing millions of patient genomes and linking them to health records, it allows Regeneron to identify novel genetic variants associated with diseases, feeding its antibody development pipeline with proprietary targets.

To overcome the scaling challenges of traditional biobanks, Regeneron is pioneering a new model. They partner with companies specializing in aggregating de-identified health records and, separately, with groups handling bio-sampling. This "uncoupled" approach allows them to link massive, independent data streams to achieve unprecedented scale.

Instead of traditional methods, Regeneron sequences millions of people to find "superhumans"—those with rare genetic mutations that protect them from diseases. By studying these individuals, they identify high-confidence drug targets that mimic these natural protections, aiming for a higher probability of success in development.

While public AI models are powerful, they risk becoming commodities when trained on the same public data. Regeneron's strategy is to create a durable advantage by training AI models on its unique dataset of millions of genomes, proteomes, and linked health records to deeply understand human biology.

To scale its database from millions to tens of millions, Regeneron is moving beyond bespoke global studies. The new model involves large-scale partnerships with health systems and consumer data sources, using privacy-preserving tokenization to securely link genetic data with vast electronic health records.

The primary bottleneck in drug development isn't creating therapies but identifying the right targets. Regeneron built its massive genetics database to find rare, protective genetic mutations in humans, effectively de-risking the target identification process and aiming to improve the industry's low success rate.

When expanding genomic studies into developing countries, Regeneron's biggest challenge is not acquiring biospecimens. The primary bottleneck is the lack of digitized electronic health records. Researchers often rely on incomplete, hard-copy medical records, hindering the ability to link genetic data with rich clinical information.

Regeneron Genetics Center's edge in AI drug discovery comes not just from its massive database, but from 14 years of interpreting high-quality, multimodal data (genomics linked to health records). This deep understanding is crucial for training reliable AI models and deriving accurate biological insights, a lesson for all life science data platforms.

Regeneron Will Use "Tokenization" to Link Disparate Datasets and Scale Genomics | RiffOn