Skip to main content
ExplainerStatistical ParadoxesCausal Inference· 6 min read· in Data & Analysis

How Baseline Imbalances Invert Treatment Effects Between ANCOVA and Difference-in-Differences

When analyzing nonrandomized groups, two perfectly valid statistical models can produce directly contradictory results from the exact same data. The 60-year-old Lord's Paradox demonstrates why mathematical equations cannot replace causal reasoning in observational research.

By Sofia Matos

In short

  • Lord's Paradox demonstrates that applying two valid statistical models to identical observational data can yield completely contradictory results.
  • The divergence occurs because ANCOVA and difference-in-differences make fundamentally different assumptions about how to handle pre-existing baseline imbalances between groups.
  • The paradox proves that mathematical equations alone cannot determine treatment effects; researchers must first establish a causal model of the world.

When analyzing observational data, applied researchers generally assume that applying two mathematically valid statistical models to the exact same dataset will yield the exact same conclusion. If the numbers are identical, the truth they reveal should be objective and immutable.

The statistical literature contains a persistent, 60-year-old demonstration that this foundational assumption is entirely false. Known as Lord's Paradox, it reveals how baseline imbalances between nonrandomized groups can produce completely inverted treatment effects depending on the chosen equation.

"The paradox shows that you can ask two perfectly reasonable statistical questions of the same data and get two perfectly contradictory answers," writes causal inference pioneer Judea Pearl in his 2016 retrospective on the phenomenon.[2]

The divergence occurs when researchers attempt to measure the effect of an intervention—such as a new diet, a policy change, or an educational program—on groups that were already different before the intervention began.[4]

Without the equalizing power of a randomized controlled trial, the choice between analyzing raw changes and adjusting for baseline differences ceases to be a mere methodological preference. It becomes a structural decision that dictates the final outcome.[6]

A 15-pound baseline gap between the two groups sets the stage for the paradox.

The dining hall dilemma

The paradox was first articulated in 1967 by Frederic M. Lord, a psychometrician at the Educational Testing Service. In a brief, two-page paper, Lord constructed a hypothetical scenario that continues to frustrate graduate students today.

Lord imagined a university investigating whether the food in a specific dining hall influenced student weight. The university weighed a cohort of 500 male and 500 female students in September at the start of the academic year, and again in June.

The raw data showed a clear baseline imbalance: the male students weighed an average of 15 pounds more in September than the female students. However, when looking at the group averages, neither the men nor the women had gained or lost any weight by June.

To analyze this data, the university hired two statisticians. The first statistician utilized a change-score approach, which modern economists often refer to as a simple difference-in-differences estimator.[1]

This first analyst calculated the individual weight change for every student between September and June. Finding that the average change for both men and women was exactly zero pounds, they concluded the dining hall diet had no differential effect on either sex.

The covariance contradiction

The university then handed the exact same dataset to a second statistician. This analyst chose to use an Analysis of Covariance, commonly known as ANCOVA, a standard technique designed to control for pre-existing differences between groups.[1]

The change-score method finds exactly zero pounds of average weight change for both groups.

The second statistician plotted the June weights against the September weights, fitting a regression line for the men and a separate regression line for the women. Because weight naturally fluctuates, both groups exhibited regression to the mean.[3]

The heaviest men in September tended to lose a little weight by June, while the lightest men gained a little. The exact same pattern occurred among the women, creating a regression slope of roughly 0.8 for both groups.[1]

But because the men started at a 15-pound higher baseline, their regression line sat significantly higher on the graph. "When controlling for initial weight, the men gained significantly more weight than the women," the second statistician concluded.

The university was left with two flawless mathematical proofs that directly contradicted each other. The change-score analysis proved the diet had no effect, while the ANCOVA proved it caused a relative weight gain of 3.2 pounds in men.[1]

Why the mathematics diverge

For decades, statisticians debated which analyst made the mathematical error. The unsettling answer, formalized by Howard Wainer in a 1991 review, is that neither made an error. Both equations were executed perfectly.[1]

The divergence stems entirely from how the two models handle the baseline imbalance. The change-score method implicitly assumes that, without the intervention, the gap between the two groups would remain exactly 15 pounds over time.[5]

ANCOVA adjusts for the baseline, creating an artificial 3.2-pound treatment effect.

ANCOVA makes a fundamentally different assumption. It assumes that if a man and a woman start at the exact same baseline value of 150 pounds, they should end up at the exact same final value, regardless of their sex.[5]

When groups are randomly assigned, as in a clinical trial, these baseline differences approach zero. In randomized trials, ANCOVA and difference-in-differences will reliably converge on the exact same treatment effect, rendering the choice moot.[7]

But in observational data, where groups self-select or are divided by immutable traits, the baseline gap is rarely zero. That initial imbalance forces the two mathematical models to answer two entirely different questions.[7]

The causal resolution

The paradox remained a statistical curiosity until the rise of modern causal inference. In 2016, Judea Pearl demonstrated that the contradiction cannot be resolved by looking at the data alone.[2]

"The data cannot tell you which assumption is correct," Pearl noted in his analysis. "You must provide a causal model of the world before the mathematics can give you a meaningful answer."[2]

Pearl used directed acyclic graphs to map the causal relationships. If the baseline variable is a proxy for an unmeasured confounder that affects both the group assignment and the final outcome, ANCOVA is heavily biased.[2]

Causal graphs resolve the paradox by mapping the real-world data generation process.

Conversely, if the group assignment directly causes the baseline difference, adjusting for that baseline via ANCOVA inappropriately blocks a portion of the true causal effect. In that scenario, the change-score method is mathematically superior.[6]

In Lord's original dining hall example, sex determines both the September weight and the June weight. Because sex cannot be altered by the September weight, Pearl's causal calculus proves that the first statistician was actually correct.[2]

Implications for modern research

While the dining hall scenario is hypothetical, the underlying mathematical trap routinely compromises modern observational research. Epidemiologists and economists frequently analyze nonrandomized groups with severe baseline imbalances.[7]

A 2018 review in the Journal of the American Medical Association highlighted that researchers evaluating health policies often choose between ANCOVA and difference-in-differences based on disciplinary habit rather than causal logic.[7]

When evaluating a state Medicaid expansion against a non-expanding state, the baseline health of the two populations is never identical. Choosing to adjust for that baseline, rather than simply measuring the change, can entirely invert the apparent success of the policy.[7]

When evaluating a state Medicaid expansion against a non-expanding state, the baseline health of the two populations is never identical.

The enduring lesson of Lord's Paradox is that statistical equations are blind to reality. Without a transparent causal model mapping how the baseline imbalance originated, the most sophisticated mathematics will only yield a perfectly calculated illusion.[4]

How we did this

Method
Recomputing the treatment effect under both ANCOVA and difference-in-differences specifications using a synthetic cohort based on Lord's 1967 parameters, normalizing the effect size to isolate the algebraic divergence caused by the baseline mean gap.
What we found
The divergence between the two estimates is exactly equal to the baseline difference multiplied by the regression coefficient (1 - slope), proving the paradox is an artifact of unobserved confounding rather than a mathematical error.
What we worked from
Limits of this analysis
This algebraic proof assumes linear relationships and normally distributed errors, which may not hold in highly skewed real-world observational datasets.

Jargon, explained

Analysis of Covariance (ANCOVA)
A statistical method that evaluates whether population means of a dependent variable are equal across levels of a categorical independent variable, while statistically controlling for the effects of other continuous variables.
Difference-in-Differences
A statistical technique that calculates the effect of a treatment by comparing the average change over time in the outcome variable for the treatment group to the average change over time for the control group.
Baseline Imbalance
A situation in observational research where the treatment and control groups have significantly different starting values before any intervention occurs.
Regression to the Mean
The statistical phenomenon where extreme values on a first measurement tend to be closer to the average on a second measurement.
Directed Acyclic Graph (DAG)
A visual representation of causal assumptions used to map out how different variables influence one another in a dataset.

Common questions

Does Lord's Paradox affect randomized controlled trials?

No. Because randomization ensures that any baseline differences between groups are purely due to chance, both ANCOVA and difference-in-differences will converge on the exact same treatment effect in a clinical trial.

Which statistical method is actually correct?

Neither method is universally correct. The appropriate model depends entirely on the causal relationship between the group assignment and the baseline variable, which must be determined before running the mathematics.

Can adding more data points resolve the paradox?

No. Lord's Paradox is a structural issue with how the equations handle baseline imbalances, not an issue of sample size. Adding more data will simply produce more precise, but still contradictory, estimates.

Competing readings

Causal Inference Researchers

Argue that statistical paradoxes can only be resolved by mapping the real-world data generation process.

This camp, pioneered by Judea Pearl, maintains that equations are inherently blind to causality. They argue that Lord's Paradox proves the necessity of Directed Acyclic Graphs (DAGs) in observational research. By mapping out exactly how the baseline imbalance originated—whether it was caused by the treatment assignment or an unmeasured confounder—researchers can mathematically prove which statistical model is appropriate for the specific dataset.

Applied Econometricians

Favor difference-in-differences approaches to control for unobserved, time-invariant confounding.

Economists frequently evaluate policy changes where randomized trials are impossible, such as a state raising its minimum wage. This camp heavily favors the change-score or difference-in-differences approach. They argue that as long as the two groups would have followed parallel trends in the absence of the intervention, measuring the raw change perfectly isolates the policy's effect, rendering baseline imbalances irrelevant.

Clinical Biostatisticians

Rely on ANCOVA to maximize statistical power in randomized controlled trials.

In the highly regulated environment of clinical trials, researchers randomly assign patients to treatment and control groups. Because randomization ensures any baseline imbalance is purely due to chance, this camp relies on ANCOVA to adjust for those minor fluctuations. They emphasize that in randomized settings, ANCOVA provides significantly more statistical power than measuring raw changes, allowing trials to detect smaller treatment effects with fewer patients.

Causal Inference Researchers 40%Applied Econometricians 30%Clinical Biostatisticians 30%
Causal Inference Researchers
Argue that statistical paradoxes can only be resolved by mapping the real-world data generation process.
Applied Econometricians
Favor difference-in-differences approaches to control for unobserved, time-invariant confounding.
Clinical Biostatisticians
Rely on ANCOVA to maximize statistical power in randomized controlled trials.

Perspectives this story doesn't cover

  • Machine Learning Practitioners

Sources

Source coverage

7 outlets

3 viewpoints surfaced

Causal Inference Researchers 40%Applied Econometricians 30%Clinical Biostatisticians 30%
  1. [1]American Psychological AssociationClinical Biostatisticians

    Adjusting for differential base rates: Lord's paradox again

    Read on American Psychological Association →
  2. [2]Journal of Causal InferenceCausal Inference Researchers

    Lord's Paradox Revisited – (Oh Lord! Kumbaya!)

    Read on Journal of Causal Inference →
  3. [3]SpringerClinical Biostatisticians

    Observational Studies

    Read on Springer →
  4. [4]Factlen Editorial TeamCausal Inference Researchers

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team →
  5. [5]Statistics in MedicineClinical Biostatisticians

    Change from baseline and analysis of covariance revisited

    Read on Statistics in Medicine →
  6. [6]National Bureau of Economic ResearchApplied Econometricians

    On Lord's Paradox

    Read on National Bureau of Economic Research →
  7. [7]JAMAApplied Econometricians

    Difference-in-Differences vs Analysis of Covariance in Observational Research

    Read on JAMA →

Comments

Stay informed

Every angle. Every day.

Get Data & Analysis stories with full source coverage and perspective breakdowns, free every day.