Skip to main content
ExplainerSurvival AnalysisKaplan-Meier· 7 min read· in Data & Analysis

How the Independent Censoring Assumption Artificially Inflates Absolute Risk in Medical Research

Standard survival models mathematically treat patients who die of other causes as if they remain at risk for the primary disease. This structural artifact systematically overestimates the probability of medical events, potentially driving overtreatment in older populations.

By Sofia Matos

In short

  • Standard survival models mathematically assume patients who die of other causes remain at risk for the primary disease, artificially inflating absolute risk estimates.
  • In older populations, this independent censoring assumption can overestimate the probability of needing medical interventions by up to 55%, potentially driving overtreatment.
  • Modern statistical guidelines require researchers to use Cumulative Incidence Functions when competing events affect more than 10% of the study population.

In clinical research, a fundamental methodological disagreement centers on how to count patients who die from something other than the disease being studied. Some trial designers treat these competing events as independent censoring, mathematically assuming the patient simply left the study.[1]

Statisticians argue this assumption creates a mathematical fiction that distorts medical guidelines. By treating a patient who died of a heart attack as if they could still theoretically develop cancer, standard survival models apply risk rates to a population of ghost patients.[4][6]

This phenomenon, known as a competing risk, occurs whenever the occurrence of one mutually exclusive event prevents the occurrence of another. In older populations, where patients are at risk for multiple conditions simultaneously, ignoring this dynamic systematically overestimates the absolute risk of the primary disease.[2]

The standard tool for time-to-event data is the Kaplan-Meier estimator, which calculates the probability of surviving past a certain time point. When researchers want to know the risk of an event happening, they frequently use the complement of this survival function.[3]

The mathematical divergence between standard estimates and actual cumulative incidence.

The Mechanics of the Risk Pool

The mathematical flaw emerges in how the Kaplan-Meier product-limit estimator handles the denominator of the risk pool. When a patient experiences a competing event, the estimator censors them, removing them from the denominator for future time points.[5]

However, the model assumes this censoring is non-informative, meaning the censored patient's future risk would have been identical to those who remained in the study. "By treating deaths as censored observations, the Kaplan-Meier method assumes the risk of revision is independent of the risk of death," researchers note.[1]

This assumption is biologically impossible when the competing event is death. A patient who dies from a stroke is no longer at risk of dying from breast cancer, yet the Kaplan-Meier math assumes they are still theoretically susceptible if the stroke had not occurred.[4]

The resulting distortion is not merely theoretical; it compounds over time. Because the denominator shrinks while the model assumes the censored patients would have continued experiencing the primary event at the average rate, the cumulative probability of the primary event is artificially inflated.[1]

Clinical Consequences of Inflated Risk

The magnitude of this overestimation can dramatically alter clinical guidelines and healthcare policy. In studies examining the need for revision surgery after hip or knee arthroplasty, the Kaplan-Meier method overestimated the cumulative incidence of revision by an aggregate of 55%.[1]

The bias becomes particularly severe in older demographics. For patients aged 65 to 74 years, where the incidence of natural mortality outpaces the need for joint revision, the standard method overestimated the risk of revision by 45% over a 20-year follow-up period.[1]

In older populations, standard models can overestimate the need for joint revision by up to 45%.

Similar distortions plague cardiovascular research. A landmark analysis found that standard Cox proportional hazard models classified 18% of elderly subjects as being at high risk for coronary heart disease, whereas models accounting for competing risks classified only 8% of the same cohort as high risk.[3]

This discrepancy directly influences patient care. Surgical decisions and pharmaceutical interventions are guided by these formal risk predictions. Ignoring competing risks can lead to overtreatment, exposing patients to the side effects of aggressive therapies they are statistically unlikely to ever need.[1][6]

When a patient is informed they have a 30.7% chance of a major adverse cardiovascular event over ten years, they may opt for aggressive statin therapy. However, when competing non-cardiovascular deaths are properly accounted for, that risk drops to 23.2%, fundamentally altering the risk-benefit calculus of the intervention.[2]

Mathematical Solutions and Alternatives

To correct this structural artifact, statisticians advocate for the Aalen-Johansen estimator, a non-parametric method that calculates the Cumulative Incidence Function. This approach explicitly accounts for competing causal pathways rather than treating them as missing data.[2]

Under the Aalen-Johansen method, patients who experience competing events are permanently classified as no longer at risk for the primary event. Consequently, the cumulative incidence function is lowered proportionally to the occurrence of the competing events, providing a mathematically sound absolute risk.[2]

For regression analysis—where researchers want to understand how different variables affect risk—the Fine-Gray subdistribution hazard model offers a robust alternative. This model directly targets the cumulative incidence function, offering a clinically interpretable measure of absolute risk that standard Cox models cannot provide.[2]

How independent censoring creates a mathematically impossible 'ghost' population.

The mathematical differences are stark when analyzing composite outcomes. "The sum of the Kaplan-Meier estimates of the incidence of each individual outcome will exceed the Kaplan-Meier estimate of the incidence of the composite outcome defined as any of the event types," cardiovascular researchers warn.[3]

Etiology Versus Absolute Risk

Despite the clear advantages of competing risk models for predicting absolute risk, the choice of methodology remains contested depending on the research question. Epidemiologists studying the pure biological cause of a disease often prefer cause-specific hazard models.[2]

Cause-specific models estimate the instantaneous risk of a specific event type while still treating competing events as censored. This approach isolates the biological effect of a risk factor, such as how smoking affects lung cancer rates, independent of whether the patient might first die of heart disease.[2]

However, using cause-specific models to communicate absolute risk to patients or policymakers is a critical error. A drug might double the biological hazard of a rare cancer, but if the patient's absolute risk of dying from a competing cardiovascular event is overwhelmingly higher, the practical risk remains negligible.[4]

To bridge this gap, modern statistical guidelines recommend reporting both metrics. Presenting the cause-specific hazard alongside the Fine-Gray subdistribution hazard enables a complete understanding of both the biological mechanism and the actual real-world probability of the outcome.[3]

This dual-reporting approach ensures transparency across the entire research lifecycle. It allows regulatory bodies to evaluate the pure efficacy of a pharmaceutical compound while simultaneously giving frontline physicians the pragmatic data they need to advise individual patients on their actual prognosis.[6]

Overestimation accelerates rapidly once competing events exceed 10% of the study population.

Thresholds for Methodological Change

Statisticians have established clear thresholds for when competing risk analysis becomes mandatory. As a general rule of thumb, researchers must abandon standard Kaplan-Meier estimates when the absolute percentage of competing events in the study population exceeds 10%.[1]

The requirement also triggers when the proportion of patients experiencing the competing event is equal to or greater than those experiencing the outcome of interest. In younger cohorts, where competing mortality is low, the Kaplan-Meier overestimation might be as small as 3%, making the standard method acceptable.[1]

The interpretation of the data must also shift. A 1999 foundational paper by Gooley et al. established that standard survival estimates can only be interpreted as a hypothetical probability, assuming the primary event's risk would remain unchanged if competing risks were somehow eradicated.[6]

This hypothetical population, immune to all other causes of death, does not exist in reality. For healthcare systems planning hospital capacity or insurers modeling lifetime costs, designing policy around this mathematically immortal cohort guarantees inefficient resource allocation.[4]

The Future of Survival Analysis

As medical datasets grow larger and more complex, the methodologies for handling competing risks are evolving. In 2026, researchers introduced advanced deep learning frameworks, such as the DH model, designed specifically to handle survival analysis with competing events.[2]

Unlike traditional parametric models, these neural networks do not assume a specific form for the time-to-event distribution. They learn the non-linear relationships between patient covariates and multiple competing risks directly from the data, offering unprecedented precision in personalized risk prediction.[2]

Illustration: New deep learning frameworks are automating the complex calibration required for competing risks.

Ensuring these complex models remain accurate requires novel calibration metrics. A January 2026 preprint highlighted that existing calibration measures are unsuited for the competing-risk setting, proposing new post-hoc recalibration methods that maintain discrimination while ensuring the probabilities sum correctly across all possible outcomes.[5]

The transition away from independent censoring assumptions represents a maturation of medical statistics. By acknowledging that human biology involves multiple competing pathways, researchers can provide patients with risk assessments grounded in reality rather than mathematical abstraction.[6]

How we did this

Method
Mathematical comparison of the risk-set denominator adjustments between the Kaplan-Meier product-limit estimator and the Aalen-Johansen cumulative incidence function in the presence of a mutually exclusive competing event.
What we found
By mathematically comparing the risk-set denominators, we demonstrate that Kaplan-Meier's overestimation is not a sampling error but a structural artifact: it applies the primary event's hazard rate to a 'ghost' population of patients who have already experienced a mutually exclusive competing event, thereby artificially inflating the absolute risk.
What we worked from
  • Kaplan-Meier independent censoring assumption: Treats competing events as non-informative censored data, assuming the risk of the primary event remains independent of the competing event. — National Institutes of Health
  • Aalen-Johansen cumulative incidence adjustment: Lowers the cumulative incidence function by permanently removing patients who experience competing events from the at-risk pool. — Taylor & Francis Online

Key terms

Independent Censoring
The statistical assumption that a patient leaving a study has the exact same future risk of the event as a patient who remains in the study.
Kaplan-Meier Estimator
A standard statistical method used to calculate the probability of surviving past a certain time point, which struggles when multiple mutually exclusive outcomes exist.
Cumulative Incidence Function (CIF)
A calculation method that accurately estimates the absolute risk of a specific event by permanently removing patients who experience competing events from the risk pool.
Fine-Gray Model
A regression model designed specifically for competing risks, allowing researchers to see how different variables affect the actual real-world probability of an event.
Cause-Specific Hazard
The instantaneous biological risk of a specific event occurring, calculated by artificially isolating it from all other competing causes of death.

Frequently asked

What exactly is a competing risk in medical research?

A competing risk is any event that makes it impossible for the primary event of interest to occur. The most common example is death from another cause; a patient who dies of a heart attack can no longer develop the cancer being studied.

Why does the Kaplan-Meier method overestimate risk?

It mathematically assumes that patients who experience a competing event simply dropped out of the study but remain at risk. By applying the ongoing hazard rate to this 'ghost' population that cannot actually experience the event, it inflates the final probability.

When should researchers use competing risk analysis?

Statistical guidelines dictate using these methods when the absolute percentage of competing events exceeds 10%, or when the number of competing events equals or exceeds the number of primary events.

Does this mean older medical studies are wrong?

Not necessarily wrong, but potentially misleading regarding absolute risk. Older studies using standard methods accurately calculated the biological hazard rate, but their final percentage estimates for how many patients would actually experience the event in the real world were likely inflated.

Viewpoints in depth

Clinical Methodologists

Advocate for the universal adoption of Aalen-Johansen and Fine-Gray models to prevent overtreatment.

This camp argues that providing patients with inflated absolute risk numbers violates informed consent. When a patient is told they have a 20% chance of needing a revision surgery, they make decisions based on that number. If their actual risk is 10% because they are highly likely to experience a competing mortality event first, the standard Kaplan-Meier estimate has misled them into potentially unnecessary anxiety or prophylactic interventions.

Etiological Researchers

Prefer cause-specific hazard models to isolate and understand pure biological mechanisms.

Researchers focused on the biological origins of disease maintain that treating competing events as censored is necessary for specific questions. If the goal is to understand whether a new chemical exposure increases the biological hazard of liver cancer, the fact that exposed subjects might die of lung cancer first is irrelevant to the chemical's hepatotoxicity. For these scientists, the hypothetical "pure" risk is exactly what they are trying to measure.

Public Health Planners

Emphasize absolute risk accuracy for resource allocation and hospital capacity planning.

Health economists and hospital administrators rely on cumulative incidence functions to model actual future demand. If a regional health authority uses standard survival models to predict the number of joint revisions required over the next decade, they will over-allocate budget and surgical theater time. This camp insists that predictive models must reflect the reality that patients who die of competing causes will never utilize those healthcare resources.

Clinical Methodologists 40%Etiological Researchers 30%Public Health Planners 30%
Clinical Methodologists
Advocate for Aalen-Johansen and Fine-Gray models to prevent overtreatment.
Etiological Researchers
Prefer cause-specific hazard models to isolate and understand pure biological mechanisms.
Public Health Planners
Emphasize absolute risk accuracy for resource allocation and hospital capacity planning.

Perspectives this story doesn't cover

  • Patients receiving inflated risk estimates
  • Pharmaceutical companies utilizing standard models for trial endpoints

Sources

Source coverage

6 outlets

3 viewpoints surfaced

Clinical Methodologists 40%Etiological Researchers 30%Public Health Planners 30%
  1. [1]National Institutes of HealthClinical Methodologists

    Overestimation of the Kaplan-Meier method in the presence of competing risks

    Read on National Institutes of Health →
  2. [2]Taylor & Francis Online

    Competing risks in clinical and epidemiological studies

    Read on Taylor & Francis Online →
  3. [3]American Heart AssociationEtiological Researchers

    Estimating the incidence of an event in the presence of competing risks

    Read on American Heart Association →
  4. [4]Columbia University Mailman School of Public HealthPublic Health Planners

    Competing Risk Analysis Overview

    Read on Columbia University Mailman School of Public Health →
  5. [5]arXiv

    Calibration in Survival Analysis with Competing Risks

    Read on arXiv →
  6. [6]Factlen Editorial TeamClinical Methodologists

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team →

Comments

Stay informed

Every angle. Every day.

Get Data & Analysis stories with full source coverage and perspective breakdowns, free every day.