Skip to main content
ExplainerCausal InferenceExplainer· 6 min read· in Data & Analysis

How the Parallel Trends Assumption Validates the Counterfactual in Difference-in-Differences Estimation

Difference-in-differences estimation relies on the untestable assumption that treated and control groups would have evolved identically without intervention. Recent methodological breakthroughs have exposed the statistical fragility of traditional pre-trend tests, forcing economists to adopt new frameworks that explicitly bound counterfactual uncertainty.

By Nicolas Laurent

Methodological Reformers 45%Classical Practitioners 35%Causal Inference Skeptics 20%
Methodological Reformers
Argue that standard pre-tests are statistically underpowered and advocate for new bounding and group-time estimators.
Classical Practitioners
Rely on foundational DiD designs and visual pre-trend checks to evaluate policy impacts.
Causal Inference Skeptics
Highlight that the parallel trends assumption is fundamentally untestable and urge caution in observational claims.

Perspectives this story doesn't cover

  • Policymakers relying on DiD estimates
  • Data journalists interpreting causal studies

In a randomized controlled trial, the counterfactual is guaranteed by the flip of a coin: because treatment is assigned randomly, the control group perfectly represents what would have happened to the treated group. Difference-in-differences estimation attempts to achieve the same causal certainty using observational data, but it differs in one structural respect: it does not assume the two groups are identical. Instead, it assumes that the gap between them is constant over time. This mathematical bridge is known as the parallel trends assumption, and it is the single load-bearing pillar of modern empirical economics.[6]

The fundamental problem of causal inference is that researchers can never observe the same unit both treated and untreated at the exact same time. To measure the impact of a policy, an analyst must construct a counterfactual—a mathematically rigorous estimate of what would have occurred if the policy had never been enacted. The difference-in-differences framework solves this by tracking a treated group and an untreated control group across two time periods: before the intervention and after.[6]

The canonical demonstration of this architecture is David Card and Alan Krueger's 1994 study on the minimum wage, which helped launch the quasi-experimental revolution in labor economics. On April 1, 1992, New Jersey raised its state minimum wage from $4.25 per hour to $5.05 per hour. This represented a 19 percent increase. Meanwhile, the minimum wage in neighboring Pennsylvania remained unchanged at $4.25, providing a natural control group just across the Delaware River.[1]

Card and Krueger surveyed 410 fast-food restaurants across both states, collecting data in February and March of 1992 before the law took effect, and again in November and December of 1992. A naive observational study might have simply looked at New Jersey's employment numbers before and after the wage hike. However, if the broader regional economy was entering a recession or a boom, that simple before-and-after comparison would absorb those macroeconomic shifts and misattribute them to the minimum wage policy.[1]

Card and Krueger's 1994 study utilized the Delaware River as a geographic boundary for a natural experiment.

By subtracting the change in Pennsylvania's employment from the change in New Jersey's employment, the researchers isolated the specific impact of the wage increase. They found a 13 percent relative increase in employment in the treated state. As they concluded in their findings, "the increase in New Jersey's minimum wage probably had no effect on total employment in New Jersey's fast-food industry, and possibly had a small positive effect."[1]

This elegant subtraction only works if one critical condition holds true: the parallel trends assumption. The assumption dictates that, in the absence of the minimum wage increase, New Jersey's fast-food employment would have evolved on the exact same trajectory as Pennsylvania's. It is a strict requirement that the unobserved potential outcomes for the treated group move in lockstep with the observed outcomes of the control group.[5]

As a 2025 econometric analysis notes, "Parallel trend is the one thing everyone 'checks' in difference-in-differences, but almost nobody can prove." The assumption is fundamentally untestable because it makes a claim about a reality that does not exist. Researchers cannot observe New Jersey's post-1992 employment data without the 1992 wage increase, meaning the true counterfactual remains permanently hidden.[3]

Researchers cannot observe New Jersey's post-1992 employment data without the 1992 wage increase, meaning the true counterfactual remains permanently hidden.

To build confidence in this unobservable metric, economists traditionally relied on pre-trends testing. By plotting the data for several years prior to the intervention, researchers could visually and statistically verify whether the treated and control groups were moving in parallel before the policy took effect. If a standard statistical test failed to reject the null hypothesis of zero difference between the pre-treatment trends, the parallel trends assumption was considered validated, and the causal estimate was published.[6]

In 2022, Harvard economist Jonathan Roth published a methodological critique that upended this standard practice. Roth demonstrated that conventional pre-trends tests suffer from a severe lack of statistical power, meaning they routinely fail to detect meaningful divergences between the two groups. Because the tests are underpowered, a researcher might observe parallel lines in the pre-period even when the underlying economic conditions are actively drifting apart.[2]

The true counterfactual is permanently hidden, forcing researchers to rely on the trajectory of the control group.

The mathematical reality of this power deficit is stark. As Carlos Mendez's 2026 synthesis of Roth's work explains, "A test with 50 observations per group may require a violation three times larger than the treatment effect to reject the null at 5% significance." This means that a confounding variable could be skewing the data by an amount 300 percent larger than the actual policy impact, and the standard pre-test would still give the researcher a green light to proceed.[2]

Conditioning an analysis on passing this flawed pre-test introduces what econometricians call pre-test bias. The estimates that survive the screening process are often distorted, creating a false sense of security regarding the causal claim. The realization that the field's primary diagnostic tool was mathematically fragile forced a rapid evolution in how difference-in-differences models are estimated and defended.[2]

The complexity compounds when policies are not implemented on a single date, but rather roll out across different regions at different times—a structure known as staggered adoption. For decades, researchers analyzed these rollouts using two-way fixed effects regressions. However, a landmark 2021 paper by Brantly Callaway and Pedro H.C. Sant'Anna revealed that this standard regression technique breaks down under staggered timing, sometimes even assigning negative weights to newly treated units by improperly using already-treated units as controls.[4]

The Callaway-Sant'Anna estimator resolved this by computing group-time average treatment effects, isolating clean comparisons between specific cohorts and never-treated control groups. By restricting the analysis to these uncontaminated pairs, the estimator prevents the mathematical pathologies of two-way fixed effects from corrupting the causal estimate, ensuring that the parallel trends assumption is applied only where it is logically coherent.[4]

To address the pre-test power problem identified by Roth, the field has shifted away from binary pass-fail testing and toward sensitivity analysis. In 2023, Ashesh Rambachan and Jonathan Roth introduced the HonestDiD framework, which allows researchers to explicitly bound the uncertainty. Instead of claiming that parallel trends hold perfectly, the framework calculates how large a violation would need to be to invalidate the study's conclusions.[2]

Conventional pre-trend tests often lack the statistical power to detect economically meaningful confounding drift.

This transition from assumption to quantification marks a maturation in empirical economics. Rather than discarding the difference-in-differences framework because its core assumption is untestable, modern methodologists have built mathematical boundaries around the uncertainty. By explicitly defining the limits of the counterfactual, researchers can now state exactly how much confounding drift their causal estimates can withstand before the signal is lost to the noise.[6]

Unsettled ground

  • How often published difference-in-differences studies from the past two decades rely on parallel trends that were actively violated but masked by underpowered pre-tests.
  • Whether the newer bounding methods like HonestDiD will become universally mandated by top economics journals, or if standard two-way fixed effects will persist in applied research.
  • The exact degree to which unobservable macroeconomic shocks differentially affected New Jersey and Pennsylvania in the months following the 1992 minimum wage increase.
410
Fast-food restaurants surveyed in 1992
19%
New Jersey minimum wage increase
13%
Relative employment increase
3x
Violation size needed to reject null (N=50)
5%
Standard significance threshold

Sources

Source coverage

6 outlets

3 viewpoints surfaced

Methodological Reformers 45%Classical Practitioners 35%Causal Inference Skeptics 20%
  1. [1]Atticus LiClassical Practitioners

    What Card And Krueger 1994 Actually Tested

    Read on Atticus Li
  2. [2]Carlos MendezMethodological Reformers

    Difference-in-Differences with Multiple Time Periods

    Read on Carlos Mendez
  3. [3]MediumCausal Inference Skeptics

    What the parallel trends assumption actually says

    Read on Medium
  4. [4]MetricGateMethodological Reformers

    Staggered DiD: Callaway-Sant'Anna Estimator

    Read on MetricGate
  5. [5]arXivMethodological Reformers

    Difference-in-Differences Identification Through the Lens of Selection

    Read on arXiv
  6. [6]Factlen Editorial TeamCausal Inference Skeptics

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team

Comments

Stay informed

Every angle. Every day.

Get Data & Analysis stories with full source coverage and perspective breakdowns delivered to your inbox.