Skip to main content
ExplainerCausal InferenceMethodology Explainer· 5 min read· in Data & Analysis

Evidence Pack: The Accuracy and Limits of Propensity Score Matching in Observational Research

For decades, researchers claimed that matching subjects on a single propensity score could mimic a randomized trial and eliminate selection bias. However, mathematical proofs reveal that propensity score matching can actually increase imbalance, inefficiency, and model dependence compared to other matching methods.

By Harper Lane

Methodological Critics 40%Applied Researchers 30%Causal Inference Theorists 30%
Methodological Critics
Argue that PSM's approximation of complete randomization actively increases imbalance and model dependence in finite samples.
Applied Researchers
Value propensity score matching for its simplicity and its ability to reduce high-dimensional data into a single, easily matched score.
Causal Inference Theorists
Emphasize that including bias-amplifying variables in the propensity model structurally worsens the estimates.

Perspectives this story doesn't cover

  • Peer Reviewers
  • Journal Editors

Key points

  • Propensity score matching (PSM) has been used in over 47,300 studies to simulate randomized trials.
  • Mathematical proofs from 2019 show that PSM often increases imbalance, inefficiency, and bias.
  • The 'PSM Paradox' occurs because pruning distant matches on the score increases the distance between actual covariates.
  • Including all available variables in the model can inadvertently amplify confounding bias.
  • Alternative methods like Mahalanobis Distance Matching balance covariates directly without dimensionality reduction.
47,300+
Articles using PSM by 2019
36 years
Gap between introduction and refutation
1
Dimensions in a propensity score

"If you measure enough baseline characteristics, you can compress them into a single score and perfectly mimic a randomized controlled trial." This claim, foundational to observational research since statisticians Paul Rosenbaum and Donald Rubin introduced it in 1983, asserts that "adjustment for the scalar propensity score is sufficient to remove bias due to all observed covariates" [1]. By calculating the conditional probability that a subject receives a treatment based on their background traits, researchers believed they could pair treated and untreated subjects to eliminate selection bias entirely [1].[1]

For 36 years, the method dominated applied statistics. By 2019, propensity score matching (PSM) had been used or referenced in over 47,300 scholarly articles across economics, epidemiology, medicine, and political science [4]. As Peter Austin detailed in his highly cited 2011 methodological review, the propensity score acts as a balancing mechanism: conditional on that single scalar value, the distribution of observed baseline covariates is assumed to be similar between the treated and untreated groups [3]. This allowed observational studies to simulate the conditions of a randomized trial, driving major policy and medical decisions when actual experiments were impossible or unethical [2, 3].[2][3][4]

But the mathematical consensus was operating on a massive blind spot. In 2019, political methodologists Gary King and Richard Nielsen published a landmark refutation in the journal Political Analysis, demonstrating that the method is fundamentally flawed when applied to real-world data. Through rigorous mathematical proofs and simulated datasets, they showed that PSM "often accomplishes the opposite of its intended goal—thus increasing imbalance, inefficiency, model dependence, and bias" [4]. The very tool designed to make observational data more reliable was actively distorting the results of thousands of studies [4, 6].[4][6]

The failure stems from what King and Nielsen termed the "PSM Paradox." When researchers use PSM, they typically prune the dataset, discarding control subjects that do not closely match treated subjects on the propensity score [4]. Initially, this pruning helps remove the most extreme outliers. But beyond a certain threshold, pruning the worst matches actually increases the distance between the remaining pairs across the original covariates [4, 6]. Instead of tightening the comparison, the algorithm begins matching subjects who look less and less alike in reality [4].[4][6]

The PSM Paradox: Pruning data beyond a certain threshold actively increases covariate imbalance.

This paradox occurs because PSM attempts to approximate a completely randomized experiment, rather than a fully blocked randomized experiment [4]. In a completely randomized trial with a finite sample size, random chance guarantees that some covariates will be unbalanced—just as flipping a coin 10 times rarely yields exactly five heads. By compressing all variables into a single dimension, PSM becomes mathematically blind to that residual imbalance, effectively shuffling the data in ways that reintroduce the exact bias it was meant to eliminate [4, 6].[4][6]

This paradox occurs because PSM attempts to approximate a completely randomized experiment, rather than a fully blocked randomized experiment [4].

Two patients can easily arrive at the exact same 45% propensity score through entirely different clinical paths. One might have a high score due to advanced age, elevated blood pressure, and a history of smoking, while the other reaches 45% due to low income, geographic location, and a specific genetic marker. Matching them solely on the 45% score ignores the massive, multidimensional imbalance in their actual clinical profiles, pairing subjects who have no business being compared in a medical trial [4, 6].[4][6]

Furthermore, the standard practice of maximizing the propensity model's predictive power by including every available pre-treatment variable introduces severe structural dangers. In 2010, computer scientist Judea Pearl demonstrated that conditioning on certain variables—specifically instrumental variables that influence treatment selection but have no direct effect on the outcome—"tends to amplify confounding bias in the analysis of causal effects" [5]. Throwing every variable into the model does not just add noise; it actively weaponizes the unobserved confounders [5, 6].[5][6]

Dimensionality loss: Two subjects can share identical propensity scores while possessing entirely different background traits.

Pearl's structural causal models proved that adding these bias-amplifying variables to a propensity score model can introduce new bias where none existed before [5]. This mathematical reality directly contradicts the prevailing advice in applied research, which often encourages researchers to include all available covariates in the logistic regression used to generate the score to ensure nothing is missed [2, 5]. By blindly feeding the algorithm more data, researchers were inadvertently amplifying the exact selection effects they were trying to control [5, 6].[2][5][6]

The evidence points toward alternative matching methods that do not rely on dimensionality reduction. Techniques like Mahalanobis Distance Matching (MDM) and Coarsened Exact Matching (CEM) approximate fully blocked experiments, ensuring that matched pairs are actually similar across all measured dimensions [4]. Unlike PSM, these methods monotonically decrease imbalance as more distant matches are pruned [4, 6]. If a researcher drops the worst match using MDM, the remaining dataset is mathematically guaranteed to be more balanced than it was before [4].[4][6]

The reliance on PSM highlights a dangerous lag between asymptotic statistical theory and applied research. While Rosenbaum and Rubin's 1983 proofs hold true in infinite samples where random matching is harmless, applying them to finite datasets triggers the exact random matching penalty that King and Nielsen exposed [1, 4]. The evidence suggests that researchers must abandon the single-score illusion and balance the actual covariates directly [6]. Continuing to use propensity score matching in observational research is no longer a defensible methodological choice [4, 6].[1][4][6]

How we got here

  1. 1983

    Paul Rosenbaum and Donald Rubin introduce Propensity Score Matching in Biometrika, proving it balances covariates in infinite samples.

  2. 2010

    Judea Pearl demonstrates that conditioning on instrumental variables amplifies confounding bias, challenging the practice of including all covariates.

  3. 2011

    Peter Austin publishes a highly cited review cementing PSM as a standard tool for observational research.

  4. 2019

    Gary King and Richard Nielsen publish mathematical proofs showing PSM increases imbalance and bias in finite datasets.

What we don’t know

  • Exactly how many of the 47,300+ published studies relying on PSM contain fundamentally reversed or distorted conclusions due to the PSM Paradox.
  • Whether applied fields like epidemiology and economics will universally transition to Coarsened Exact Matching (CEM) or Mahalanobis Distance Matching (MDM) in the near future.
  • How to perfectly identify and exclude all bias-amplifying instrumental variables from observational datasets before matching.

Sources

Source coverage

6 outlets

3 viewpoints surfaced

Methodological Critics 40%Applied Researchers 30%Causal Inference Theorists 30%
  1. [1]Biometrika

    The central role of the propensity score in observational studies for causal effects

    Read on Biometrika
  2. [2]Statistical ScienceApplied Researchers

    Matching Methods for Causal Inference: A Review and a Look Forward

    Read on Statistical Science
  3. [3]Multivariate Behavioral ResearchApplied Researchers

    An Introduction to Propensity Score Methods for Reducing the Effects of Confounding in Observational Studies

    Read on Multivariate Behavioral Research
  4. [4]Political AnalysisMethodological Critics

    Why Propensity Scores Should Not Be Used for Matching

    Read on Political Analysis
  5. [5]ArXivCausal Inference Theorists

    On a Class of Bias-Amplifying Variables that Endanger Effect Estimates

    Read on ArXiv
  6. [6]Factlen Editorial TeamMethodological Critics

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team

Comments

Stay informed

Every angle. Every day.

Get Data & Analysis stories with full source coverage and perspective breakdowns delivered to your inbox.