Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

A cardiovascular risk algorithm from the UK Biobank, trained on a primarily Caucasian dataset, was found by Human Longevity Inc. to be inaccurate for Asian populations. This highlights a critical and dangerous data bias in widely used genomic models.

Wei-Wu He: Craig Venter’s Legacy and the Future of Human Longevity thumbnail

Wei-Wu He: Craig Venter’s Legacy and the Future of Human Longevity

Behind the Breakthroughs·11 days ago

Related Insights

Many genetic tests for personalized nutrition are validated on narrow populations, like European Caucasians. These genetic markers often have zero predictive power when applied to other ethnic groups, such as those of West African descent, making their recommendations highly unreliable for a diverse user base.

Much of the data on global health conditions is not collected locally in developing countries. Instead, it is extrapolated from data on wealthy, Western populations, leading to biased models and a flawed understanding of disease prevalence worldwide.

The burgeoning field of polygenic risk scores is dangerously unregulated, with some well-capitalized companies selling products that are 'no better than chance.' The key differentiator is rigorous, public validation of their predictive models, especially across ancestries, a step many firms skip.

One-third of Regeneron's 3 million sequenced genomes are of non-European ancestry. This is a deliberate scientific choice, not just an ethical one, to create a richer dataset. It avoids the inherent scientific limitations of homogenous data, leading to more powerful and broadly applicable biological discoveries.

To combat the data equity problem where wearable users are often affluent, ŌURA actively partners with research organizations. By donating thousands of rings for studies on specific groups (e.g., women with diabetes), they acquire diverse datasets essential for building inclusive and accurate health algorithms.

A lack of representation in genomic data has direct clinical consequences. A deep understanding of European genetics and a poor understanding of other groups has already manifested in less precise medical treatments for non-European populations, undermining the core promise of precision medicine.

DNA Complete's model of providing raw genomic risk scores tied to individual scientific papers, without context or curation, can be dangerously misleading. A user might see a low-risk result for a disease that is irrelevant to their ethnicity, highlighting the critical need for proper data interpretation in consumer health.

Regeneron's focus on diverse populations is a core research strategy. Key discoveries, like the PCSK9 heart disease mutation, were only possible because they were significantly more common in African Americans. This proves that diverse genomic data unlocks unique and powerful therapeutic targets that would otherwise be missed.

Advanced health tech faces a fundamental problem: a lack of baseline data for what constitutes "optimal" health versus merely "not diseased." We can identify deficiencies but lack robust, ethnically diverse databases defining what "great" health looks like, creating a "North Star" problem for personalization algorithms.

Leading longevity research relies on datasets like the UK Biobank, which predominantly features wealthy, Western individuals. This creates a critical validation gap, meaning AI-driven biomarkers may be inaccurate or ineffective for entire populations, such as South Asians, hindering equitable healthcare advances.