The All of Us research biobank recently announced a massive new data release – making it the largest health database in the world that allows researchers to access both donated genomic information as well as participants’ electronic medical records (EMRs). That combination can allow researchers to associate the genomic variations that we all have with health outcomes (e.g. developing cancer) – while also controlling things that can impact our health like behavior (e.g., smoking) or environment (e.g., living in a city polluted with smog). These kinds of non-biological factors can protect us from, or accelerate, the impact of genomic variation, which makes them critical to know for research.
While the biobank has existed since 2018, the enhanced ability to connect donated samples and genomic sequencing with EMR data is new. Previously, the biobank mostly relied on participants to connect and share their own EMR, or participating hospitals agreeing to share it on their behalf. It is now utilizing the eHealth Exchange health information network for clinical data sharing to facilitate the process.
What makes All of Us particularly special, and therefore this news even more impactful on genomic science, is its focus on recruiting participants historically under-represented in research. Roughly 80% of All of Us participants are from such communities. This is in stark contrast to global genomic databases where almost 80% of participants in genome-wide association studies are of European descent. Even in government-supported databases, like the NHGRI-EBI GWAS Catalogue, only 2.4% of participants are of African ancestry.
A lack of diversity in genomic research is not just bad for communities under-represented in these databases – it is bad for everyone. Yes, as a bioethicist, I would of course argue that it is bad to live in an inequitable society at the existential level, but as a scientist, I also know that the lack of data diversity can limit health advances for everyone. For example, the high genetic variation within populations of African ancestry can make participants uniquely situated to contribute to scientific advancements. That 2.4% of participants with African ancestry in the NHGRI-EBI GWAS catalogue have contributed to 7% of findings regarding the association of genetic variation with outcomes.
Representation in genomic research also matters because not only can different genomic variations sometimes be associated with ancestry, but the impact of health disparities can also be associated with self-identified race and ethnicity due to factors including institutionalized racism. But to be able to identify improved health outcomes generalizable to, or targeted toward, historically excluded research populations, researchers need access to representative data.
As law and ethics faculty at the University of Michigan Medical School, my work focuses on the use of health data for medical research and ways to increase access. I was also the principal investigator of a 5-year grant from NHGRI focusing on how academic genetic researchers chose data for their work, and how to improve the system to encourage and support the use of representative health data. What we found is critical to fully understanding the potential impact of All of Us.
When we interviewed genetic researchers and asked them how they chose databases for their work, they all talked about the importance of large numbers of participants but not one brought up looking for genomic diversity. Although when pressed, most agreed it was important. But when we surveyed hundreds of researchers, we found that the vast majority reported wanting to work with data from non-European ancestries.
But why would researchers report that they want to work with representative data and admit that they don’t bother to look for them? Most likely because representative data do not generally exist in the first place.
It’s very time-consuming and complicated for researchers to collect their own data from participants from scratch, meaning, in some cases, actually walking from patient waiting room to waiting room asking people to spit in a tube for research. Doing that will also just reflect the differences in demographics of people who are sitting in that waiting room (including that they likely have insurance) – overall often making it a painful, and not always fruitful, process.
Another possible solution is for researchers to pull data from several different databases to combine non-Euro-centric data. But here they can also face expensive and time-consuming tasks related to “harmonizing” those data to standardize health comparisons. In addition, many top-tier journals are more likely to publish findings from larger databases. If researchers feel pressure to work quickly and publish in high impact journals for promotion (which, it turns out, they do), they are more likely to turn to existing databases. If they turn to existing databases, their research can only be as representative as the data already contained in them. You see how the problem becomes cyclical.
While some have argued that supporting researchers with different kinds of priorities would rectify these problems, we found that a lack of access to representative data was a systemic issue. We did not find an association between researcher demographics and ancestral populations represented in their publications. Data availability was consistently the major driver of use, meaning that data diversity cannot be resolved just by individual researchers alone. It is a collective action problem.
There remain challenges to using the All of Us biobank at scale. For example, just because the majority of data housed in All of Us are representative in comparison to those historically included in research, they are sometimes problematically different from each other. Genomic research often requires large numbers of homogeneous participants to compare variation and outcomes. Researchers might still have to resort to harmonizing data across different databases to have enough comparison points.
In addition, All of Us generally has a more comprehensive participant consent and community oversight process than other biobanks. Other studies have found that differences in consent rates to genomic research can be associated with race and ethnicity. The All of Us approach likely contributed to its ability to build a representative database in the first place. But specific consent preferences can also limit researchers’ ability to use and harmonize data to build larger cohorts if databases have different consent rules.
Overall, All of Us’ recent announcement is excellent news for genetic researchers and patients alike. But our research has found that challenges remain in translating representative data into representative research. Without dedicated investment in harmonization infrastructure to combine databases, even representative databases might fail to accomplish broader health goals.
Kayte Spector-Bagdady, JD, MBe is the Wantz Professor of Bioethics at the University of Michigan Medical School