Skip to main content
Predictive AIEvidence PackAug 20, 2026, 6:56 AM· 3 min read· in data analysis

Harvard and Dana-Farber AI Tool Predicts Risk of 348 Diseases Using EHR and Genetic Data

A new Bayesian generative model called ALADYNOULLI integrates longitudinal health records and polygenic risk scores to forecast disease trajectories up to a decade in advance. Validated on over 683,000 patient records, the tool outperforms standard clinical calculators while revealing hidden biological signatures.

By Harper Lane

Clinical Researchers 40%Preventive Medicine Advocates 35%Medical AI Skeptics 25%
Clinical Researchers
Focus on the model's ability to uncover hidden biological mechanisms.
Preventive Medicine Advocates
Emphasize the potential for early intervention and proactive care.
Medical AI Skeptics
Highlight the need for prospective validation and the risks of algorithmic bias.

Why this matters

By predicting how interconnected diseases evolve over a lifetime rather than treating them in isolation, this tool could shift medicine from reactive treatment to proactive, personalized prevention for conditions ranging from heart disease to metastatic cancer.

Researchers from Harvard Medical School, Dana-Farber Cancer Institute, and Mass General Brigham have built an artificial intelligence model capable of predicting a patient's risk for 348 different diseases years before symptoms appear. The algorithm, detailed in a recent Nature publication, achieves this by combining a patient's electronic health record (EHR) history with their underlying genetic profile.[1][2]

The model, named ALADYNOULLI, departs from traditional diagnostic tools that evaluate one disease at a time. Instead, it uses a Bayesian generative framework to map out how interconnected conditions evolve over a patient's lifetime. This approach allows the system to continuously update its risk assessments as new medical data is added, effectively improving its predictive power as the patient ages.[1][2][5]

To build and validate the tool, the research team utilized data from over 683,000 individuals across three major databases: the UK Biobank, Mass General Brigham, and the All of Us research program. The dataset included up to 52 years of follow-up information, providing a deep longitudinal view of how diseases manifest and progress.[1][5]

The model was trained on massive longitudinal datasets spanning up to 52 years.

By analyzing this massive dataset, ALADYNOULLI identified 21 "latent disease signatures." These are hidden biological patterns that capture underlying disease processes, such as metabolic-inflammatory pathways, which drive clusters of seemingly unrelated conditions.[1][5]

The integration of germline genetics is a critical component of the model's success. By incorporating polygenic risk scores alongside clinical data, the algorithm uncovered 151 genome-wide significant loci. Notably, this included cardiovascular associations that single-trait analyses had previously missed, demonstrating the value of a holistic, multi-disease approach.[1][5]

The integration of germline genetics is a critical component of the model's success.

For clinical prediction, the tool has shown remarkable accuracy. It outperformed established risk scores across 28 conditions over both one-year and ten-year horizons. Specifically, it beat the Pooled Cohort Equation (PCE) and PREVENT models for cardiovascular risk, as well as the GAIL model for breast cancer risk.[1][2]

ALADYNOULLI outperformed established clinical risk scores for cardiovascular and breast cancer predictions.

In practice, the AI captured clinically expected sequences—such as hypercholesterolemia preceding a myocardial infarction, or primary cancers preceding metastatic disease—but with added nuance. The model revealed that the cardiovascular signature rose more rapidly before a heart attack in early-onset cases compared to late-onset cases, suggesting different biological mechanisms are at play despite the same final diagnosis.[5]

At Dana-Farber Cancer Institute, researchers are already exploring how the tool can be applied to oncology. The team is using ALADYNOULLI to identify different patterns of cancer metastases, attempting to predict whether a tumor is likely to spread early, late, widely, or narrowly.[3][4]

This capability could help identify previously unrecognized subtypes of cancer. By understanding the biological reasons why patients progress along different metastatic trajectories, oncologists hope to identify more beneficial, targeted therapeutics.[3]

Researchers at Dana-Farber are using the tool to predict different trajectories of cancer metastases.

Despite these promising results, the evidence remains retrospective. The tool has currently only been tested on past patient data. Its real-world clinical utility, and whether these early warnings actually change doctor behavior or improve patient outcomes, has yet to be prospectively validated in active care settings.[2][5][6]

Furthermore, while the model uses an explicit likelihood formulation to account for selection bias through inverse probability weighting, AI tools trained on historical EHRs can still inherit systemic biases present in the underlying data. The researchers acknowledge that the signatures are concordant with established disease biology, but prospective clinical trials will be necessary to ensure equitable performance across diverse populations.[1][6]

The next phase of development involves expanding the model's signatures to increase the accuracy and biological grounding of its risk predictions. If successfully implemented in clinical practice, ALADYNOULLI could represent a significant step toward personalized, predictive medicine, allowing doctors to intervene years before pathology becomes clinically apparent.[2][3][6]

Viewpoints in depth

Clinical Researchers

Focus on the model's ability to uncover hidden biological mechanisms.

For clinical researchers, the primary value of ALADYNOULLI lies in its ability to reveal biological subtypes within traditional diagnostic categories. By identifying 21 latent disease signatures, the model demonstrates that two patients with the same diagnosis often have different underlying signature profiles. This multidimensional view of disease progression allows researchers to understand why certain patients respond differently to the same treatment, potentially paving the way for more targeted therapeutics and a deeper understanding of disease pathogenesis.

Preventive Medicine Advocates

Emphasize the potential for early intervention and proactive care.

Advocates for preventive medicine view this technology as a paradigm shift from reactive to proactive healthcare. By accurately predicting the likelihood of conditions like Alzheimer's or cardiovascular disease years before symptoms manifest, the tool provides a critical window for early intervention. This head start allows patients to make lifestyle changes, undergo more frequent monitoring, or start preventive treatments when they are most effective, fundamentally altering the trajectory of their long-term health.

Medical AI Skeptics

Highlight the need for prospective validation and the risks of algorithmic bias.

Skeptics and bioethicists caution that while retrospective results are impressive, the true test of any medical AI is prospective clinical validation. They point out that algorithms trained on historical electronic health records can inadvertently learn and perpetuate systemic biases present in the data. Furthermore, there is concern about how such predictive tools will be integrated into clinical workflows without overwhelming physicians with false positives or unactionable risk alerts, stressing that the tool must prove it can actually improve patient outcomes in real-world settings.

What we don’t know

  • Whether the model's predictions will actually change physician behavior or improve patient outcomes in real-world clinical settings.
  • How the algorithm performs across highly diverse, underrepresented populations not fully captured in the primary biobanks.
  • The exact timeline for when this tool might be integrated into standard electronic health record systems for everyday clinical use.

Sources

Source coverage

6 outlets

3 viewpoints surfaced

Clinical Researchers 40%Preventive Medicine Advocates 35%Medical AI Skeptics 25%
  1. [1]NatureClinical Researchers

    A Bayesian framework for longitudinal EHR and genetic discovery

    Read on Nature
  2. [2]Harvard Medical SchoolClinical Researchers

    New AI Tool Predicts Risk of More Than 300 Diseases With Existing Patient Data

    Read on Harvard Medical School
  3. [3]Dana-Farber Cancer InstituteClinical Researchers

    Algorithm predicts cancer risk from EHR data: Dana-Farber, Mass General Brigham

    Read on Dana-Farber Cancer Institute
  4. [4]Becker's OncologyPreventive Medicine Advocates

    Algorithm predicts cancer risk from EHR data: Dana-Farber, Mass General Brigham - Becker's Oncology

    Read on Becker's Oncology
  5. [5]News MedicalPreventive Medicine Advocates

    A Bayesian framework for longitudinal EHR and genetic discovery

    Read on News Medical
  6. [6]Factlen Editorial TeamMedical AI Skeptics

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team

Comments

Stay informed

Every angle. Every day.

Get data analysis stories with full source coverage and perspective breakdowns delivered to your inbox.