How the Parallel Trends Assumption Validates the Counterfactual in Difference-in-Differences Estimation
Difference-in-differences estimation relies on the untestable assumption that treated and control groups would have evolved identically without intervention. Recent methodological breakthroughs have exposed the statistical fragility of traditional pre-trend tests, forcing economists to adopt new frameworks that explicitly bound counterfactual uncertainty.
- Methodological Reformers
- Argue that standard pre-tests are statistically underpowered and advocate for new bounding and group-time estimators.
- Classical Practitioners
- Rely on foundational DiD designs and visual pre-trend checks to evaluate policy impacts.
- Causal Inference Skeptics
- Highlight that the parallel trends assumption is fundamentally untestable and urge caution in observational claims.
Perspectives this story doesn't cover
- Policymakers relying on DiD estimates
- Data journalists interpreting causal studies
In a randomized controlled trial, the counterfactual is guaranteed by the flip of a coin: because treatment is assigned randomly, the control group perfectly represents what would have happened to the treated group. Difference-in-differences estimation attempts to achieve the same causal certainty using observational data, but it differs in one structural respect: it does not assume the two groups are identical. Instead, it assumes that the gap between them is constant over time. This mathematical bridge is known as the parallel trends assumption, and it is the single load-bearing pillar of modern empirical economics.[6]
The fundamental problem of causal inference is that researchers can never observe the same unit both treated and untreated at the exact same time. To measure the impact of a policy, an analyst must construct a counterfactual—a mathematically rigorous estimate of what would have occurred if the policy had never been enacted. The difference-in-differences framework solves this by tracking a treated group and an untreated control group across two time periods: before the intervention and after.[6]
The canonical demonstration of this architecture is David Card and Alan Krueger's 1994 study on the minimum wage, which helped launch the quasi-experimental revolution in labor economics. On April 1, 1992, New Jersey raised its state minimum wage from $4.25 per hour to $5.05 per hour. This represented a 19 percent increase. Meanwhile, the minimum wage in neighboring Pennsylvania remained unchanged at $4.25, providing a natural control group just across the Delaware River.[1]
Card and Krueger surveyed 410 fast-food restaurants across both states, collecting data in February and March of 1992 before the law took effect, and again in November and December of 1992. A naive observational study might have simply looked at New Jersey's employment numbers before and after the wage hike. However, if the broader regional economy was entering a recession or a boom, that simple before-and-after comparison would absorb those macroeconomic shifts and misattribute them to the minimum wage policy.[1]
By subtracting the change in Pennsylvania's employment from the change in New Jersey's employment, the researchers isolated the specific impact of the wage increase. They found a 13 percent relative increase in employment in the treated state. As they concluded in their findings, "the increase in New Jersey's minimum wage probably had no effect on total employment in New Jersey's fast-food industry, and possibly had a small positive effect."[1]
This elegant subtraction only works if one critical condition holds true: the parallel trends assumption. The assumption dictates that, in the absence of the minimum wage increase, New Jersey's fast-food employment would have evolved on the exact same trajectory as Pennsylvania's. It is a strict requirement that the unobserved potential outcomes for the treated group move in lockstep with the observed outcomes of the control group.[5]
As a 2025 econometric analysis notes, "Parallel trend is the one thing everyone 'checks' in difference-in-differences, but almost nobody can prove." The assumption is fundamentally untestable because it makes a claim about a reality that does not exist. Researchers cannot observe New Jersey's post-1992 employment data without the 1992 wage increase, meaning the true counterfactual remains permanently hidden.[3]
Researchers cannot observe New Jersey's post-1992 employment data without the 1992 wage increase, meaning the true counterfactual remains permanently hidden.
To build confidence in this unobservable metric, economists traditionally relied on pre-trends testing. By plotting the data for several years prior to the intervention, researchers could visually and statistically verify whether the treated and control groups were moving in parallel before the policy took effect. If a standard statistical test failed to reject the null hypothesis of zero difference between the pre-treatment trends, the parallel trends assumption was considered validated, and the causal estimate was published.[6]
In 2022, Harvard economist Jonathan Roth published a methodological critique that upended this standard practice. Roth demonstrated that conventional pre-trends tests suffer from a severe lack of statistical power, meaning they routinely fail to detect meaningful divergences between the two groups. Because the tests are underpowered, a researcher might observe parallel lines in the pre-period even when the underlying economic conditions are actively drifting apart.[2]
The mathematical reality of this power deficit is stark. As Carlos Mendez's 2026 synthesis of Roth's work explains, "A test with 50 observations per group may require a violation three times larger than the treatment effect to reject the null at 5% significance." This means that a confounding variable could be skewing the data by an amount 300 percent larger than the actual policy impact, and the standard pre-test would still give the researcher a green light to proceed.[2]
Conditioning an analysis on passing this flawed pre-test introduces what econometricians call pre-test bias. The estimates that survive the screening process are often distorted, creating a false sense of security regarding the causal claim. The realization that the field's primary diagnostic tool was mathematically fragile forced a rapid evolution in how difference-in-differences models are estimated and defended.[2]
The complexity compounds when policies are not implemented on a single date, but rather roll out across different regions at different times—a structure known as staggered adoption. For decades, researchers analyzed these rollouts using two-way fixed effects regressions. However, a landmark 2021 paper by Brantly Callaway and Pedro H.C. Sant'Anna revealed that this standard regression technique breaks down under staggered timing, sometimes even assigning negative weights to newly treated units by improperly using already-treated units as controls.[4]
The Callaway-Sant'Anna estimator resolved this by computing group-time average treatment effects, isolating clean comparisons between specific cohorts and never-treated control groups. By restricting the analysis to these uncontaminated pairs, the estimator prevents the mathematical pathologies of two-way fixed effects from corrupting the causal estimate, ensuring that the parallel trends assumption is applied only where it is logically coherent.[4]
To address the pre-test power problem identified by Roth, the field has shifted away from binary pass-fail testing and toward sensitivity analysis. In 2023, Ashesh Rambachan and Jonathan Roth introduced the HonestDiD framework, which allows researchers to explicitly bound the uncertainty. Instead of claiming that parallel trends hold perfectly, the framework calculates how large a violation would need to be to invalidate the study's conclusions.[2]
This transition from assumption to quantification marks a maturation in empirical economics. Rather than discarding the difference-in-differences framework because its core assumption is untestable, modern methodologists have built mathematical boundaries around the uncertainty. By explicitly defining the limits of the counterfactual, researchers can now state exactly how much confounding drift their causal estimates can withstand before the signal is lost to the noise.[6]
Unsettled ground
- How often published difference-in-differences studies from the past two decades rely on parallel trends that were actively violated but masked by underpowered pre-tests.
- Whether the newer bounding methods like HonestDiD will become universally mandated by top economics journals, or if standard two-way fixed effects will persist in applied research.
- The exact degree to which unobservable macroeconomic shocks differentially affected New Jersey and Pennsylvania in the months following the 1992 minimum wage increase.
Sources
[1]Atticus LiClassical PractitionersWhat Card And Krueger 1994 Actually Tested
Read on Atticus Li →
[2]Carlos MendezMethodological ReformersDifference-in-Differences with Multiple Time Periods
Read on Carlos Mendez →
[3]MediumCausal Inference SkepticsWhat the parallel trends assumption actually says
Read on Medium →
[4]MetricGateMethodological ReformersStaggered DiD: Callaway-Sant'Anna Estimator
Read on MetricGate →
[5]arXivMethodological ReformersDifference-in-Differences Identification Through the Lens of Selection
Read on arXiv →
[6]Factlen Editorial TeamCausal Inference SkepticsSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
Comments
More in Data & Analysis
See all →Algorithmic Ranking
How the Elo Rating System Adjusts Player Scores Based on the Logistic Function of Expected Win Probability
3 sources
Meta-Analysis
How the I² Statistic Quantifies the Percentage of Variation in a Meta-Analysis Due to Heterogeneity
6 sources
Sequential Testing
How the Alpha-Spending Function Prevents False Positives When Continuously Monitoring A/B Tests
4 sources
Wealth Inequality
CBS Poll Finds 73% of Americans Believe the Income Gap Between the Richest and Middle Class is Increasing
5 sources
Every angle. Every day.
Get Data & Analysis stories with full source coverage and perspective breakdowns delivered to your inbox.




