Factlen ResearchPrecision MedicineEvidence PackJul 1, 2026, 12:44 PM· 6 min read· #2 of 2 in health

NIH's All of Us Program Unveils World's Largest Integrated Genomic and Clinical Dataset for Precision Medicine

The National Institutes of Health has released an unprecedented dataset linking over 535,000 whole genomes to 482,000 electronic health records, accelerating AI-driven precision medicine.

By Factlen Editorial Team

Precision Medicine Researchers 40%Health Equity Advocates 35%Medical Informaticians 25%
Precision Medicine Researchers
Focus on the unprecedented scale and multiomics integration as the necessary fuel for AI-driven drug discovery.
Health Equity Advocates
Emphasize the 86 percent representation of historically underrepresented communities as a vital correction to historical biases.
Medical Informaticians
Highlight the technical achievements of harmonizing disparate EHRs while cautioning about the inherent noise in real-world clinical data.

What's not represented

  • · Primary Care Physicians

Why this matters

By combining deep genetic data with real-world clinical outcomes from highly diverse populations, this dataset provides the exact infrastructure needed to train the next generation of medical AI and develop targeted treatments that work for everyone, not just a narrow demographic.

Key points

  • The NIH has released the world's largest integrated genomic and electronic health record database, encompassing over 747,000 participants.
  • The dataset successfully links 535,000 whole genome sequences with 482,000 clinical records, providing unprecedented longitudinal health data.
  • A record 86% of participants come from communities historically underrepresented in biomedical research, correcting long-standing biases.
  • The program has entered the multiomics era, adding proteomics and RNA sequencing data alongside 96 million NLP-derived clinical concepts.
747,000
Total participants with available data
535,000
Whole genome sequences
482,000
Linked electronic health records
86%
Participants from underrepresented communities
96 million
NLP-derived concept codes extracted

The paradox at the heart of precision medicine is that tailoring treatments to a single individual requires uncovering patterns across massive populations. For decades, medical research has been constrained by siloed data: genomic databases rarely connected to day-to-day clinical outcomes, and clinical trials rarely reflected the full diversity of the population. On June 30, 2026, the National Institutes of Health (NIH) fundamentally altered that landscape. The agency's "All of Us" Research Program issued its most expansive data release to date, officially establishing the platform as the world's largest integrated genomic and electronic health record (EHR) database.

The sheer scale of the newly available dataset represents a landmark achievement in biomedical infrastructure. According to the NIH, the platform now houses data from more than 747,000 participants across all fifty states and territories. Within that cohort, researchers have successfully linked more than 535,000 whole genome sequences to nearly 482,000 electronic health records. This combination of deep genetic sequencing and broad clinical history provides a longitudinal view of human health that is currently unmatched by any other research initiative globally.

The primary claim driving the program's expansion is that integrating real-world clinical data with genomics is essential for the next generation of medical discovery. The evidence supporting this approach is robust. By utilizing the Observational Medical Outcomes Partnership (OMOP) Common Data Model, the program harmonizes disparate data from over fifty different healthcare organizations and sixteen different EHR vendors into a single, standardized format. This allows scientists to query the database for complex intersections of genetics, lifestyle, and disease outcomes without having to clean and standardize the raw data themselves.[3]

The unprecedented scale of the newly released All of Us dataset.
The unprecedented scale of the newly released All of Us dataset.

A significant breakthrough in this latest release is the program's novel approach to acquiring electronic health records. Historically, gathering EHRs required complex, institution-by-institution data sharing agreements. As reported by STAT News, the NIH has now successfully implemented participant-mediated EHR submissions and health information exchange (HIE) networks to fill data gaps. This strategic shift drove a 22 percent growth in EHR data in a single release, adding records for tens of thousands of participants who receive care outside of major academic medical centers.[1]

Beyond structured data like diagnosis codes and lab results, the evidence pack now includes an unprecedented depth of unstructured clinical phenotyping. For the first time, registered researchers have secure access to clinical notes. Using advanced Natural Language Processing (NLP) tools, the program has extracted 96 million concept codes from 9.5 million clinical notes across more than 99,000 participants. This allows researchers to analyze the nuanced, narrative descriptions written by physicians, capturing symptoms and social determinants of health that rarely make it into standardized billing codes.

The dataset is also officially entering the "multiomics" era, moving beyond DNA to examine how genes are expressed and regulated. The NIH release includes proteomics data—the study of proteins produced by the body—from nearly 10,000 participants, alongside RNA sequencing data from nearly 9,000 individuals. When combined with long-read whole genome sequences, this multiomic data provides a high-resolution map of cellular function, offering critical evidence for researchers developing targeted therapies for complex diseases like cancer and Alzheimer's.

To bridge the gap between clinical visits, the program relies heavily on wearable technology. The evidence base now includes the world's largest accessible Fitbit dataset, capturing active tracking data for 68,000 participants. New longitudinal sleep data tables provide granular insights into daily sleep stages, including REM, light, and deep sleep metrics. By linking continuous, objective measurements of physical activity and sleep directly to genomic profiles and EHRs, scientists can investigate how daily behaviors influence long-term health outcomes with unprecedented precision.[3]

Strategic shifts in data collection drove a 22 percent increase in linked electronic health records.
Strategic shifts in data collection drove a 22 percent increase in linked electronic health records.
To bridge the gap between clinical visits, the program relies heavily on wearable technology.

Perhaps the most scientifically significant claim of the All of Us program is its ability to correct historical biases in medical research. For generations, genomic databases and clinical trials have overwhelmingly relied on populations of European descent, limiting the applicability of findings for other groups. The evidence from the latest release demonstrates a massive shift: 86 percent of the 747,000 participants come from communities that have been historically underrepresented in biomedical research.

Furthermore, nearly half of the participants identify with a racial or ethnic minority group. NIH Director Dr. Jay Bhattacharya noted that the genomic data from participants of non-European ancestry in this release far surpasses any comparable resource in the world. This diversity is not merely a demographic achievement; it is a scientific necessity. It ensures that the algorithms, diagnostics, and therapeutics developed from this data will be effective across the entire human population, rather than a narrow subset.[4]

Despite the unprecedented scale and utility of the dataset, transparent uncertainty remains regarding data fidelity and integration challenges. Comparing medical history derived from EHRs against self-reported survey answers often reveals discrepancies. EHRs are primarily designed for billing and clinical workflow, not for research, meaning they can contain missing data, coding errors, or institutional biases. While the OMOP Common Data Model standardizes the format, it cannot entirely eliminate the underlying noise inherent in real-world clinical data.[2][3]

The introduction of NLP-derived concept codes from clinical notes also carries inherent uncertainty. While the extraction tools are highly advanced, natural language processing in medicine must navigate complex jargon, abbreviations, and context-dependent negations. Researchers utilizing this unstructured data must account for potential false positives and extraction errors when building predictive models or identifying patient cohorts.[2][4]

How AI extracts structured medical concepts from unstructured physician notes.
How AI extracts structured medical concepts from unstructured physician notes.

Similarly, the reliance on consumer-grade wearables like Fitbits introduces variability. While these devices provide valuable longitudinal data on steps and sleep, they are not clinical-grade medical instruments. Variations in device wear-time, battery life, and proprietary algorithms for calculating sleep stages mean that researchers must apply rigorous statistical controls when correlating wearable data with hard clinical outcomes.[3]

Privacy and data security represent another area of ongoing scrutiny. Assembling a population-scale dataset linking DNA, clinical notes, and daily location-tracking data creates an inherently high-value target for cyberattacks. The NIH mitigates this risk by housing the data in a highly secure, cloud-based Researcher Workbench. The data is stripped of direct identifiers, and researchers must undergo rigorous identity verification, ethics training, and institutional approval before gaining access. Furthermore, data cannot be downloaded locally; all analyses must be conducted within the secure cloud environment.[4]

The scientific payoff of this massive infrastructure investment is already materializing. The All of Us dataset has fueled more than 1,400 peer-reviewed publications by nearly 23,000 registered researchers across the globe. These studies range from identifying novel genetic variants associated with cardiovascular disease to mapping the social determinants of mental health disparities.

The program relies on hundreds of thousands of volunteers to build a more equitable future for medicine.
The program relies on hundreds of thousands of volunteers to build a more equitable future for medicine.

As artificial intelligence models become increasingly central to drug discovery and diagnostics, the demand for high-quality, multimodal training data has skyrocketed. Claims that AI can outperform human doctors run well ahead of the current clinical evidence, largely because medical AI has historically been trained on fragmented, biased datasets. The All of Us program provides exactly the kind of genomic-plus-clinical trove that science-specific AI models require to narrow that gap.[4]

By turning the research paradigm into a two-way street—returning personalized health-related DNA results to participants while aggregating their data for global science—the NIH has built a sustainable engine for discovery. As the program marches toward its goal of one million participants, it is not just building a database; it is constructing the foundational public utility for the future of precision medicine.[4]

How we got here

  1. 2015

    The Precision Medicine Initiative is launched, laying the groundwork for the All of Us Research Program.

  2. 2018

    The All of Us Research Program officially opens national enrollment to the public.

  3. 2020

    The program begins returning genetic ancestry and trait results to its participants.

  4. 2022

    The program starts returning personalized health-related DNA results, detailing disease risks and medication processing.

  5. June 2026

    The NIH releases the largest dataset to date, integrating 535,000 genomes with 482,000 EHRs and entering the multiomics era.

Viewpoints in depth

Precision Medicine Researchers

Focus on the sheer scale and the multiomics data enabling new discoveries.

For computational biologists and AI researchers, the All of Us dataset represents the holy grail of medical data. The primary bottleneck in developing predictive models for complex diseases has not been algorithm design, but the lack of massive, multimodal datasets. By linking whole genome sequences with longitudinal EHRs and wearable data, researchers argue they can finally train AI models to identify subtle, multi-variable patterns that predict disease onset years before symptoms appear. The addition of proteomics and RNA sequencing further excites this camp, as it bridges the gap between static DNA risk and active cellular dysfunction.

Health Equity Advocates

Focus on the 86% representation of underrepresented communities and how it corrects historical biases.

Public health experts and equity advocates view the demographic composition of the All of Us dataset as its most important achievement. Historically, the vast majority of genomic data used to develop targeted therapies came from individuals of European descent. This systemic bias meant that genetic tests and precision drugs were often less effective—or even inaccurate—for minority populations. By ensuring that 86 percent of its participants come from historically underrepresented groups, advocates argue the NIH is actively preventing the future of AI-driven medicine from inheriting the biases of the past.

Medical Informaticians

Focus on the technical challenges of harmonizing EHRs, OMOP standards, and the noise in NLP extraction.

Data scientists and medical informaticians celebrate the dataset's scale but caution against treating real-world clinical data as ground truth. They point out that electronic health records are fundamentally designed for billing and administrative workflows, not scientific research. Consequently, EHRs are rife with missing data, upcoding, and institutional variations. While the program's use of the OMOP Common Data Model and advanced NLP extraction tools are state-of-the-art, informaticians emphasize that researchers must rigorously account for this inherent 'noise' when drawing clinical conclusions from the database.

What we don't know

  • How accurately NLP-derived concept codes from unstructured clinical notes will map to actual patient outcomes in predictive models.
  • Whether the integration of consumer-grade wearable data will prove robust enough to serve as primary endpoints in future clinical trials.
  • How quickly the newly added multiomics data (proteomics and RNAseq) will translate into actionable, FDA-approved targeted therapies.

Key terms

Precision Medicine
An approach to disease treatment and prevention that takes into account individual variability in genes, environment, and lifestyle.
Electronic Health Record (EHR)
A digital version of a patient's paper chart, containing medical history, diagnoses, medications, and treatment plans.
Multiomics
A biological analysis approach that combines data from multiple 'omes,' such as the genome (DNA), proteome (proteins), and transcriptome (RNA).
Natural Language Processing (NLP)
A branch of artificial intelligence that helps computers understand, interpret, and manipulate human language, used here to extract data from doctor's notes.
Whole Genome Sequencing
A comprehensive laboratory process that determines the entirety of an individual's DNA sequence at a single time.

Frequently asked

What is the All of Us Research Program?

It is a landmark NIH initiative aiming to collect health data from one million diverse Americans to accelerate precision medicine and tailored treatments.

How does the program protect participant privacy?

Data is stripped of direct identifiers and housed in a highly secure, cloud-based Researcher Workbench. Researchers cannot download the data and must undergo rigorous ethics training.

Who can access this dataset?

Registered researchers who complete identity verification and institutional approval can access the data for approved scientific studies.

Why is diversity a primary focus of the program?

Historically, medical research has relied on populations of European descent. Diverse data ensures that new algorithms and treatments are effective for all demographic groups.

Sources

Source coverage

4 outlets

3 viewpoints surfaced

Precision Medicine Researchers 40%Health Equity Advocates 35%Medical Informaticians 25%
  1. [1]STAT NewsHealth Equity Advocates

    STAT+: All of Us tests a new approach to collect real-world data for research

    Read on STAT News
  2. [2]Journal of Medical Internet ResearchMedical Informaticians

    Evolution of Secondary Use of Electronic Health Records and Their Interoperability in Medical Research

    Read on Journal of Medical Internet Research
  3. [3]Journal of the American Medical Informatics AssociationMedical Informaticians

    Comparing medical history data derived from electronic health records and survey answers in the All of Us Research Program

    Read on Journal of the American Medical Informatics Association
  4. [4]Factlen Editorial TeamMedical Informaticians

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team
Stay informed

Every angle. Every day.

Get health stories with full source coverage and perspective breakdowns delivered to your inbox.