Evidence Pack: The Accuracy and Limits of Propensity Score Matching in Observational Research
For decades, researchers claimed that matching subjects on a single propensity score could mimic a randomized trial and eliminate selection bias. However, mathematical proofs reveal that propensity score matching can actually increase imbalance, inefficiency, and model dependence compared to other matching methods.
By Harper Lane
- Methodological Critics
- Argue that PSM's approximation of complete randomization actively increases imbalance and model dependence in finite samples.
- Applied Researchers
- Value propensity score matching for its simplicity and its ability to reduce high-dimensional data into a single, easily matched score.
- Causal Inference Theorists
- Emphasize that including bias-amplifying variables in the propensity model structurally worsens the estimates.
Perspectives this story doesn't cover
- Peer Reviewers
- Journal Editors
Key points
- Propensity score matching (PSM) has been used in over 47,300 studies to simulate randomized trials.
- Mathematical proofs from 2019 show that PSM often increases imbalance, inefficiency, and bias.
- The 'PSM Paradox' occurs because pruning distant matches on the score increases the distance between actual covariates.
- Including all available variables in the model can inadvertently amplify confounding bias.
- Alternative methods like Mahalanobis Distance Matching balance covariates directly without dimensionality reduction.
- 47,300+
- Articles using PSM by 2019
- 36 years
- Gap between introduction and refutation
- 1
- Dimensions in a propensity score
"If you measure enough baseline characteristics, you can compress them into a single score and perfectly mimic a randomized controlled trial." This claim, foundational to observational research since statisticians Paul Rosenbaum and Donald Rubin introduced it in 1983, asserts that "adjustment for the scalar propensity score is sufficient to remove bias due to all observed covariates" [1]. By calculating the conditional probability that a subject receives a treatment based on their background traits, researchers believed they could pair treated and untreated subjects to eliminate selection bias entirely [1].[1]
For 36 years, the method dominated applied statistics. By 2019, propensity score matching (PSM) had been used or referenced in over 47,300 scholarly articles across economics, epidemiology, medicine, and political science [4]. As Peter Austin detailed in his highly cited 2011 methodological review, the propensity score acts as a balancing mechanism: conditional on that single scalar value, the distribution of observed baseline covariates is assumed to be similar between the treated and untreated groups [3]. This allowed observational studies to simulate the conditions of a randomized trial, driving major policy and medical decisions when actual experiments were impossible or unethical [2, 3].[2][3][4]
But the mathematical consensus was operating on a massive blind spot. In 2019, political methodologists Gary King and Richard Nielsen published a landmark refutation in the journal Political Analysis, demonstrating that the method is fundamentally flawed when applied to real-world data. Through rigorous mathematical proofs and simulated datasets, they showed that PSM "often accomplishes the opposite of its intended goal—thus increasing imbalance, inefficiency, model dependence, and bias" [4]. The very tool designed to make observational data more reliable was actively distorting the results of thousands of studies [4, 6].[4][6]
The failure stems from what King and Nielsen termed the "PSM Paradox." When researchers use PSM, they typically prune the dataset, discarding control subjects that do not closely match treated subjects on the propensity score [4]. Initially, this pruning helps remove the most extreme outliers. But beyond a certain threshold, pruning the worst matches actually increases the distance between the remaining pairs across the original covariates [4, 6]. Instead of tightening the comparison, the algorithm begins matching subjects who look less and less alike in reality [4].[4][6]
This paradox occurs because PSM attempts to approximate a completely randomized experiment, rather than a fully blocked randomized experiment [4]. In a completely randomized trial with a finite sample size, random chance guarantees that some covariates will be unbalanced—just as flipping a coin 10 times rarely yields exactly five heads. By compressing all variables into a single dimension, PSM becomes mathematically blind to that residual imbalance, effectively shuffling the data in ways that reintroduce the exact bias it was meant to eliminate [4, 6].[4][6]
This paradox occurs because PSM attempts to approximate a completely randomized experiment, rather than a fully blocked randomized experiment [4].
Two patients can easily arrive at the exact same 45% propensity score through entirely different clinical paths. One might have a high score due to advanced age, elevated blood pressure, and a history of smoking, while the other reaches 45% due to low income, geographic location, and a specific genetic marker. Matching them solely on the 45% score ignores the massive, multidimensional imbalance in their actual clinical profiles, pairing subjects who have no business being compared in a medical trial [4, 6].[4][6]
Furthermore, the standard practice of maximizing the propensity model's predictive power by including every available pre-treatment variable introduces severe structural dangers. In 2010, computer scientist Judea Pearl demonstrated that conditioning on certain variables—specifically instrumental variables that influence treatment selection but have no direct effect on the outcome—"tends to amplify confounding bias in the analysis of causal effects" [5]. Throwing every variable into the model does not just add noise; it actively weaponizes the unobserved confounders [5, 6].[5][6]
Pearl's structural causal models proved that adding these bias-amplifying variables to a propensity score model can introduce new bias where none existed before [5]. This mathematical reality directly contradicts the prevailing advice in applied research, which often encourages researchers to include all available covariates in the logistic regression used to generate the score to ensure nothing is missed [2, 5]. By blindly feeding the algorithm more data, researchers were inadvertently amplifying the exact selection effects they were trying to control [5, 6].[2][5][6]
The evidence points toward alternative matching methods that do not rely on dimensionality reduction. Techniques like Mahalanobis Distance Matching (MDM) and Coarsened Exact Matching (CEM) approximate fully blocked experiments, ensuring that matched pairs are actually similar across all measured dimensions [4]. Unlike PSM, these methods monotonically decrease imbalance as more distant matches are pruned [4, 6]. If a researcher drops the worst match using MDM, the remaining dataset is mathematically guaranteed to be more balanced than it was before [4].[4][6]
The reliance on PSM highlights a dangerous lag between asymptotic statistical theory and applied research. While Rosenbaum and Rubin's 1983 proofs hold true in infinite samples where random matching is harmless, applying them to finite datasets triggers the exact random matching penalty that King and Nielsen exposed [1, 4]. The evidence suggests that researchers must abandon the single-score illusion and balance the actual covariates directly [6]. Continuing to use propensity score matching in observational research is no longer a defensible methodological choice [4, 6].[1][4][6]
How we got here
1983
Paul Rosenbaum and Donald Rubin introduce Propensity Score Matching in Biometrika, proving it balances covariates in infinite samples.
2010
Judea Pearl demonstrates that conditioning on instrumental variables amplifies confounding bias, challenging the practice of including all covariates.
2011
Peter Austin publishes a highly cited review cementing PSM as a standard tool for observational research.
2019
Gary King and Richard Nielsen publish mathematical proofs showing PSM increases imbalance and bias in finite datasets.
What we don’t know
- Exactly how many of the 47,300+ published studies relying on PSM contain fundamentally reversed or distorted conclusions due to the PSM Paradox.
- Whether applied fields like epidemiology and economics will universally transition to Coarsened Exact Matching (CEM) or Mahalanobis Distance Matching (MDM) in the near future.
- How to perfectly identify and exclude all bias-amplifying instrumental variables from observational datasets before matching.
Sources
[1]BiometrikaThe central role of the propensity score in observational studies for causal effects
Read on Biometrika →
[2]Statistical ScienceApplied ResearchersMatching Methods for Causal Inference: A Review and a Look Forward
Read on Statistical Science →
[3]Multivariate Behavioral ResearchApplied ResearchersAn Introduction to Propensity Score Methods for Reducing the Effects of Confounding in Observational Studies
Read on Multivariate Behavioral Research →
[4]Political AnalysisMethodological CriticsWhy Propensity Scores Should Not Be Used for Matching
Read on Political Analysis →
[5]ArXivCausal Inference TheoristsOn a Class of Bias-Amplifying Variables that Endanger Effect Estimates
Read on ArXiv →
[6]Factlen Editorial TeamMethodological CriticsSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
Comments
More in Data & Analysis
See all →Data Visualization
How the Width of Links in a Sankey Diagram Represents Flow Magnitude and Preserves Conservation
7 sources
Causal Inference
How the Parallel Trends Assumption Validates the Counterfactual in Difference-in-Differences Estimation
6 sources
Macroeconomic Modeling
Evidence Pack: The Accuracy of High-Frequency Alternative Data in Macroeconomic Nowcasting
5 sources
Mechanistic AI
Mechanistic Machine Learning Models Uncover Physical Laws, Shifting AI From Pattern Recognition to Active Scientific Discovery
6 sources
Every angle. Every day.
Get Data & Analysis stories with full source coverage and perspective breakdowns delivered to your inbox.




