Skip to main content
Statistical IllusionsExplainer· 4 min read· in Content Types

How Aggregation Reverses Trends in Simpson's Paradox

When data is combined across different groups, underlying trends can completely reverse, leading to false conclusions in medicine, law, and sports. By isolating confounding variables, statisticians can separate true performance from mathematical artifacts.

By Elena Castillo

Causal Inference Statisticians 40%Data Scientists 35%Policy Analysts 25%
Causal Inference Statisticians
Argue that observational data cannot be accurately analyzed without a causal model to identify confounding variables before aggregation.
Data Scientists
Focus on automated stratification and clustering techniques to detect hidden reversals in massive datasets.
Policy Analysts
Warn that aggregate statistics are frequently weaponized in legal and institutional settings to prove discrimination where none exists.

Perspectives this story doesn't cover

  • Machine Learning Engineers
  • Healthcare Administrators

Key terms

Simpson's Paradox
A phenomenon in probability and statistics where a trend appears in several different groups of data but disappears or reverses when these groups are combined.
Confounding Variable
An unmeasured third variable that influences both the supposed cause and the supposed effect, creating a spurious association.
Stratification
The process of dividing members of the population into homogeneous subgroups before analyzing the data.
Marginal Association
The overall relationship between two variables when all other variables are ignored or aggregated.
Partial Association
The relationship between two variables while holding a third confounding variable constant.

Key points

  1. Simpson's Paradox occurs when a trend in individual data groups reverses when the groups are combined.
  2. The reversal is caused by unequal sample sizes and unmeasured confounding variables.
  3. David Justice had a higher batting average than Derek Jeter every year from 1995 to 1997, but a lower overall average.
  4. A 1973 UC Berkeley admissions study showed aggregate bias against women that disappeared when stratified by department.
  5. Failing to account for the paradox in medical data can lead to recommending inferior treatments.

Simpson's Paradox occurs because unequal sample sizes across different subgroups mathematically distort the overall average when combined. A trend that is true in every individual category can completely reverse in the aggregate, creating a statistical illusion that misleads researchers and algorithms alike.

The phenomenon relies on a confounding variable—a hidden factor that influences both the independent and dependent variables. When data is pooled without accounting for this confounder, the sheer volume of observations in one subgroup overwhelms the others. The resulting aggregate metric no longer measures the actual performance or outcome; instead, it measures the distribution of the population across those subgroups. Despite the "paradox" label, there is no magic involved—only a failure to compare like with like.

The mechanics of this reversal are most visible in the batting averages of Derek Jeter and David Justice between 1995 and 1997. During those three seasons, Justice posted a higher batting average than Jeter in every single year. In 1995, Justice hit .253 to Jeter's .250; in 1996, he hit .321 to Jeter's .314; and in 1997, he hit .329 to Jeter's .291. Yet, when the three seasons are combined, Jeter emerges with the higher overall average: .300 compared to Justice's .298.

The reversal happens because the number of at-bats varied wildly for both players across the three years. Justice recorded 411 at-bats in 1995, a year when batting averages were generally low, dragging down his overall percentage. Jeter, meanwhile, recorded only 48 at-bats that year. In 1996, when both players hit exceptionally well, Jeter logged 582 at-bats while Justice, sidelined by injury, recorded only 140. Jeter's aggregate average is heavily weighted toward his best season, while Justice's is anchored to his worst.

Despite trailing in every individual season, unequal at-bat distributions gave Derek Jeter the higher aggregate average.

While sports statistics offer a clean mathematical demonstration, the paradox has triggered severe legal and institutional consequences. In the fall of 1973, the University of California, Berkeley faced a potential lawsuit over its graduate admissions data. The aggregate figures showed a stark disparity: the university admitted 44.5 percent of its 2,691 male applicants, but only 30.4 percent of its 1,835 female applicants. The sample odds ratio of 1.83 heavily suggested systemic bias against women.[1]

While sports statistics offer a clean mathematical demonstration, the paradox has triggered severe legal and institutional consequences.

However, when statisticians Peter Bickel, Eugene Hammel, and J.W. O'Connell disaggregated the data by department, the bias vanished. As the researchers noted in their 1975 publication, "Examination of the disaggregated data reveals few decision-making units that show statistically significant departures from expected frequencies of female admissions, and about as many units appear to favor women as to favor men."[1]

The confounding variable was the competitiveness of the departments themselves. Women disproportionately applied to highly competitive programs with low acceptance rates, such as English, while men applied to less competitive programs with high acceptance rates, such as mechanical engineering. The aggregate data penalized women for their choice of major, not their gender, demonstrating how raw demographic percentages can fabricate evidence of discrimination where none exists.[1]

The aggregate admissions data suggested severe bias, but departmental stratification revealed women were applying to far more competitive programs.

In medicine, ignoring Simpson's Paradox can lead to fatal treatment decisions. A landmark 1986 study published in The BMJ compared two treatments for kidney stones: traditional open surgery and a newer, less invasive procedure called percutaneous nephrolithotomy. The aggregate data showed the newer procedure was more successful, curing 83 percent of patients compared to 78 percent for open surgery.[2]

But when researchers stratified the patients by the size of their kidney stones, open surgery proved superior for both small and large stones. The illusion stemmed from how doctors assigned treatments. They routinely prescribed the highly effective open surgery for severe cases with large stones, which naturally have a lower baseline success rate. The newer procedure was reserved for mild cases with small stones, which are easier to cure.[2]

The aggregate success rate of the new procedure was inflated by the mildness of the cases it treated, not the efficacy of the intervention. If a hospital administrator had looked only at the top-line numbers, they would have mandated the inferior treatment for all patients, actively worsening patient outcomes based on mathematically sound but contextually blind data.[2]

By reserving the highly effective open surgery for the most severe cases, doctors inadvertently lowered its aggregate success rate.

As automated decision-making systems ingest massive datasets, the risk of Simpson's Paradox scales proportionally. Machine learning algorithms trained on pooled observational data will mathematically encode these reversed trends unless engineers explicitly program them to recognize confounding variables. An algorithm evaluating the 1986 kidney stone data without stratification would automatically recommend the inferior treatment, scaling a statistical error into a systemic failure.

Detecting the paradox requires moving beyond raw aggregation and applying causal inference models to observational data. Statisticians must identify potential confounders before calculating averages, ensuring that subgroups are compared on an equal basis. Until the underlying distribution of the data is mapped, an aggregate percentage remains a description of the sample, not a measurement of the truth.[3]

Frequently asked

What is a confounding variable?

A confounding variable is a hidden factor that influences both the independent and dependent variables in a study, skewing the results if not accounted for.

How do you fix Simpson's Paradox?

The paradox is resolved by stratifying the data—breaking the aggregate numbers down into their underlying subgroups to compare like with like.

Does this mean aggregate data is always wrong?

Not always, but aggregate data is highly vulnerable to distortion when the underlying subgroups have vastly different sample sizes or baseline conditions.

Sources

Source coverage

4 outlets

3 viewpoints surfaced

Causal Inference Statisticians 40%Data Scientists 35%Policy Analysts 25%
  1. [1]ScienceCausal Inference Statisticians

    Sex bias in graduate admissions: data from Berkeley

    Read on Science
  2. [2]The BMJ

    Comparison of treatment of renal calculi by open surgery, percutaneous nephrolithotomy, and extracorporeal shockwave lithotripsy

    Read on The BMJ
  3. [3]Encyclopedia BritannicaCausal Inference Statisticians

    Simpson's paradox | Definition, Examples, & Facts

    Read on Encyclopedia Britannica
  4. [4]Factlen Editorial TeamPolicy Analysts

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team

Comments

Stay informed

Every angle. Every day.

Get Content Types stories with full source coverage and perspective breakdowns delivered to your inbox.