NIH's All of Us Program Unveils World's Largest Integrated Genomic and Clinical Dataset for Precision Medicine
The National Institutes of Health has released an unprecedented dataset linking over 535,000 whole genomes to 482,000 electronic health records, accelerating AI-driven precision medicine.
By Maya Khalil
- Precision Medicine Researchers
- Focus on the unprecedented scale and multiomics integration as the necessary fuel for AI-driven drug discovery.
- Health Equity Advocates
- Emphasize the 86 percent representation of historically underrepresented communities as a vital correction to historical biases.
- Medical Informaticians
- Highlight the technical achievements of harmonizing disparate EHRs while cautioning about the inherent noise in real-world clinical data.
Perspectives this story doesn't cover
- Primary Care Physicians
The paradox at the heart of precision medicine is that tailoring treatments to a single individual requires uncovering patterns across massive populations. For decades, medical research has been constrained by siloed data: genomic databases rarely connected to day-to-day clinical outcomes, and clinical trials rarely reflected the full diversity of the population. On June 30, 2026, the National Institutes of Health (NIH) fundamentally altered that landscape. The agency's "All of Us" Research Program issued its most expansive data release to date, officially establishing the platform as the world's largest integrated genomic and electronic health record (EHR) database.
The sheer scale of the newly available dataset represents a landmark achievement in biomedical infrastructure. According to the NIH, the platform now houses data from more than 747,000 participants across all fifty states and territories. Within that cohort, researchers have successfully linked more than 535,000 whole genome sequences to nearly 482,000 electronic health records. This combination of deep genetic sequencing and broad clinical history provides a longitudinal view of human health that is currently unmatched by any other research initiative globally.
The primary claim driving the program's expansion is that integrating real-world clinical data with genomics is essential for the next generation of medical discovery. The evidence supporting this approach is robust. By utilizing the Observational Medical Outcomes Partnership (OMOP) Common Data Model, the program harmonizes disparate data from over fifty different healthcare organizations and sixteen different EHR vendors into a single, standardized format. This allows scientists to query the database for complex intersections of genetics, lifestyle, and disease outcomes without having to clean and standardize the raw data themselves.[3]
A significant breakthrough in this latest release is the program's novel approach to acquiring electronic health records. Historically, gathering EHRs required complex, institution-by-institution data sharing agreements. As reported by STAT News, the NIH has now successfully implemented participant-mediated EHR submissions and health information exchange (HIE) networks to fill data gaps. This strategic shift drove a 22 percent growth in EHR data in a single release, adding records for tens of thousands of participants who receive care outside of major academic medical centers.[1]
Beyond structured data like diagnosis codes and lab results, the evidence pack now includes an unprecedented depth of unstructured clinical phenotyping. For the first time, registered researchers have secure access to clinical notes. Using advanced Natural Language Processing (NLP) tools, the program has extracted 96 million concept codes from 9.5 million clinical notes across more than 99,000 participants. This allows researchers to analyze the nuanced, narrative descriptions written by physicians, capturing symptoms and social determinants of health that rarely make it into standardized billing codes.
The dataset is also officially entering the "multiomics" era, moving beyond DNA to examine how genes are expressed and regulated. The NIH release includes proteomics data—the study of proteins produced by the body—from nearly 10,000 participants, alongside RNA sequencing data from nearly 9,000 individuals. When combined with long-read whole genome sequences, this multiomic data provides a high-resolution map of cellular function, offering critical evidence for researchers developing targeted therapies for complex diseases like cancer and Alzheimer's.
To bridge the gap between clinical visits, the program relies heavily on wearable technology. The evidence base now includes the world's largest accessible Fitbit dataset, capturing active tracking data for 68,000 participants. New longitudinal sleep data tables provide granular insights into daily sleep stages, including REM, light, and deep sleep metrics. By linking continuous, objective measurements of physical activity and sleep directly to genomic profiles and EHRs, scientists can investigate how daily behaviors influence long-term health outcomes with unprecedented precision.[3]
To bridge the gap between clinical visits, the program relies heavily on wearable technology.
Perhaps the most scientifically significant claim of the All of Us program is its ability to correct historical biases in medical research. For generations, genomic databases and clinical trials have overwhelmingly relied on populations of European descent, limiting the applicability of findings for other groups. The evidence from the latest release demonstrates a massive shift: 86 percent of the 747,000 participants come from communities that have been historically underrepresented in biomedical research.
Furthermore, nearly half of the participants identify with a racial or ethnic minority group. NIH Director Dr. Jay Bhattacharya noted that the genomic data from participants of non-European ancestry in this release far surpasses any comparable resource in the world. This diversity is not merely a demographic achievement; it is a scientific necessity. It ensures that the algorithms, diagnostics, and therapeutics developed from this data will be effective across the entire human population, rather than a narrow subset.[4]
Despite the unprecedented scale and utility of the dataset, transparent uncertainty remains regarding data fidelity and integration challenges. Comparing medical history derived from EHRs against self-reported survey answers often reveals discrepancies. EHRs are primarily designed for billing and clinical workflow, not for research, meaning they can contain missing data, coding errors, or institutional biases. While the OMOP Common Data Model standardizes the format, it cannot entirely eliminate the underlying noise inherent in real-world clinical data.[2][3]
The introduction of NLP-derived concept codes from clinical notes also carries inherent uncertainty. While the extraction tools are highly advanced, natural language processing in medicine must navigate complex jargon, abbreviations, and context-dependent negations. Researchers utilizing this unstructured data must account for potential false positives and extraction errors when building predictive models or identifying patient cohorts.[2][4]
Similarly, the reliance on consumer-grade wearables like Fitbits introduces variability. While these devices provide valuable longitudinal data on steps and sleep, they are not clinical-grade medical instruments. Variations in device wear-time, battery life, and proprietary algorithms for calculating sleep stages mean that researchers must apply rigorous statistical controls when correlating wearable data with hard clinical outcomes.[3]
Privacy and data security represent another area of ongoing scrutiny. Assembling a population-scale dataset linking DNA, clinical notes, and daily location-tracking data creates an inherently high-value target for cyberattacks. The NIH mitigates this risk by housing the data in a highly secure, cloud-based Researcher Workbench. The data is stripped of direct identifiers, and researchers must undergo rigorous identity verification, ethics training, and institutional approval before gaining access. Furthermore, data cannot be downloaded locally; all analyses must be conducted within the secure cloud environment.[4]
The scientific payoff of this massive infrastructure investment is already materializing. The All of Us dataset has fueled more than 1,400 peer-reviewed publications by nearly 23,000 registered researchers across the globe. These studies range from identifying novel genetic variants associated with cardiovascular disease to mapping the social determinants of mental health disparities.
As artificial intelligence models become increasingly central to drug discovery and diagnostics, the demand for high-quality, multimodal training data has skyrocketed. Claims that AI can outperform human doctors run well ahead of the current clinical evidence, largely because medical AI has historically been trained on fragmented, biased datasets. The All of Us program provides exactly the kind of genomic-plus-clinical trove that science-specific AI models require to narrow that gap.[4]
By turning the research paradigm into a two-way street—returning personalized health-related DNA results to participants while aggregating their data for global science—the NIH has built a sustainable engine for discovery. As the program marches toward its goal of one million participants, it is not just building a database; it is constructing the foundational public utility for the future of precision medicine.[4]
Key points
- The NIH has released the world's largest integrated genomic and electronic health record database, encompassing over 747,000 participants.
- The dataset successfully links 535,000 whole genome sequences with 482,000 clinical records, providing unprecedented longitudinal health data.
- A record 86% of participants come from communities historically underrepresented in biomedical research, correcting long-standing biases.
- The program has entered the multiomics era, adding proteomics and RNA sequencing data alongside 96 million NLP-derived clinical concepts.
- 747,000
- Total participants with available data
- 535,000
- Whole genome sequences
- 482,000
- Linked electronic health records
- 86%
- Participants from underrepresented communities
- 96 million
- NLP-derived concept codes extracted
What we don’t know
- How accurately NLP-derived concept codes from unstructured clinical notes will map to actual patient outcomes in predictive models.
- Whether the integration of consumer-grade wearable data will prove robust enough to serve as primary endpoints in future clinical trials.
- How quickly the newly added multiomics data (proteomics and RNAseq) will translate into actionable, FDA-approved targeted therapies.
Frequently asked
What is the All of Us Research Program?
It is a landmark NIH initiative aiming to collect health data from one million diverse Americans to accelerate precision medicine and tailored treatments.
How does the program protect participant privacy?
Data is stripped of direct identifiers and housed in a highly secure, cloud-based Researcher Workbench. Researchers cannot download the data and must undergo rigorous ethics training.
Who can access this dataset?
Registered researchers who complete identity verification and institutional approval can access the data for approved scientific studies.
Why is diversity a primary focus of the program?
Historically, medical research has relied on populations of European descent. Diverse data ensures that new algorithms and treatments are effective for all demographic groups.
Sources
[1]STAT NewsHealth Equity AdvocatesSTAT+: All of Us tests a new approach to collect real-world data for research
Read on STAT News →
[2]Journal of Medical Internet ResearchMedical InformaticiansEvolution of Secondary Use of Electronic Health Records and Their Interoperability in Medical Research
Read on Journal of Medical Internet Research →
[3]Journal of the American Medical Informatics AssociationMedical InformaticiansComparing medical history data derived from electronic health records and survey answers in the All of Us Research Program
Read on Journal of the American Medical Informatics Association →
[4]Factlen Editorial TeamMedical InformaticiansSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
Comments
More in Health
See all →Exercise Physiology
How Estrogen and Progesterone Shift Exercise Fuel Selection, and Why the Luteal Phase Fat-Oxidation Advantage is Smaller Than Advertised
6 sources
Circadian Rhythms
What Actually Causes the Afternoon Energy Crash, and How to Shift the Circadian Dip
5 sources
Erectile Dysfunction
How PDE5 Inhibitors Block Enzyme Degradation to Sustain Smooth Muscle Relaxation
5 sources
Brain Energy
The Phosphocreatine Shuttle: How Creatine Monohydrate Drives ATP Regeneration and the Evidence for its Role in Brain Energy
6 sources
Every angle. Every day.
Get Health stories with full source coverage and perspective breakdowns delivered to your inbox.




