Skip to main content
ExplainerEconometricsDifference-in-Differences· 6 min read· in Data & Analysis

Already-Treated Controls Subtract Evolving Treatment Effects: Why Two-Way Fixed Effects Assign Negative Weights in Staggered Rollouts

When policies roll out at different times, standard statistical models mistakenly use early adopters as control groups for late adopters. If the policy's impact grows over time, this mathematical quirk subtracts the evolving effect, assigning negative weights that can artificially reverse the policy's true outcome.

By Sofia Matos

In short

  • Standard statistical models mistakenly use early policy adopters as control groups for late adopters during staggered rollouts.
  • If a policy's effect grows over time, subtracting the early adopter's trend mathematically assigns a negative weight to the true outcome.
  • Modern heterogeneity-robust estimators solve this by strictly forbidding the algorithm from using already-treated units as counterfactuals.

When an economist sits down to evaluate a policy that rolls out state by state, they must choose a mathematical engine to isolate the effect. The researcher decides which variables to hold constant, what timeline to measure, and how to separate the policy's impact from background noise.[6]

For twenty years, the default engine for this task was the two-way fixed effects regression. This statistical model became the undisputed standard in applied microeconomics, appearing in thousands of papers evaluating everything from minimum wage hikes to healthcare expansions.[1][3]

The model relies on a foundational logic called difference-in-differences. If California raises its minimum wage in 2018 and Nevada does not, the researcher compares the change in California's employment to the change in Nevada's employment over the exact same period.[6]

Nevada serves as the counterfactual, representing what would have happened in California had the policy never passed. By subtracting Nevada's baseline trend from California's observed outcome, the model isolates the specific effect of the wage hike.[1]

The Staggered Rollout Trap

This simple two-by-two comparison works flawlessly when a policy happens exactly once. But real-world legislation rarely follows such a clean schedule, and policies typically cascade across jurisdictions over several years.[6]

When states adopt a policy at different times, researchers face a staggered rollout. California might act in 2018, Oregon in 2019, and Washington in 2020, creating a complex matrix of treatment timing that a simple two-by-two grid cannot accommodate.[1][2]

Staggered rollouts create a complex matrix of treatment timing that standard 2x2 models cannot easily accommodate.

To handle this complexity, researchers feed the entire dataset into a two-way fixed effects regression. The algorithm mathematically averages all the available comparisons into a single, neat coefficient that represents the policy's overall average effect.[3]

The model accomplishes this by controlling for fixed differences between states and fixed shocks across time. It assumes that once these baseline differences are stripped away, every remaining variation in the data represents the pure impact of the policy.[1][3]

But in 2021, econometricians discovered a fatal flaw in how the algorithm constructs that average. The model does not just compare treated states to states that never adopted the policy; it also compares newly treated states to states that adopted the policy years earlier.[1][4]

Unpacking the Black Box

Andrew Goodman-Bacon published a landmark theorem in the Journal of Econometrics that cracked open the algorithm's black box. He proved that the single coefficient is actually a weighted average of every possible two-by-two comparison in the dataset.[1]

"The TWFE estimator is a weighted average of all two-group/two-period DiD estimators in the data," Goodman-Bacon wrote. His decomposition revealed exactly which states the algorithm was pairing together behind the scenes.[1]

Some of these pairings make intuitive sense, like comparing a state treated in 2019 against a state that is never treated. The algorithm assigns heavy mathematical weights to these clean comparisons, anchoring the final estimate in solid counterfactuals.[1][6]

Other pairings compare early adopters to late adopters before the late adopters pass the policy. Because neither state has changed its status during that specific window, this comparison also yields a valid measure of the policy's impact.[1]

The Goodman-Bacon decomposition reveals how the algorithm weights different comparisons, including the flawed use of already-treated units as controls.

The critical failure occurs in the final type of pairing. The algorithm routinely uses early adopters as the control group for late adopters, comparing Washington's 2020 policy change against California's 2020 baseline.[1][2]

The Poisoned Control Group

Using California as a control group in 2020 requires a massive mathematical assumption. The algorithm must assume that California's 2018 policy has already reached its maximum effect and that its current trajectory perfectly mirrors a world without the policy.[3][4]

In reality, economic policies almost always have dynamic treatment effects that evolve over time. A new tax credit might take three years to reach peak enrollment, meaning the policy's impact grows steadily larger with each passing year.[4][6]

If California's policy effect is still growing in 2020, its baseline trend is fundamentally poisoned. The state is experiencing a delayed surge from its own 2018 legislation, making it a terrible proxy for what Washington would have looked like without the 2020 law.[1][4]

When the algorithm subtracts California's poisoned trend from Washington's outcome, it accidentally subtracts the evolving treatment effect. The model interprets California's delayed surge as a background economic shock, and penalizes Washington's estimate accordingly.[1][3]

Clément de Chaisemartin and Xavier D'Haultfœuille quantified this disaster in the American Economic Review in 2020. They demonstrated that this subtraction mechanism can mathematically assign a negative weight to the true treatment effect of the late-adopting state.[3]

The Mathematics of Negative Weights

A negative weight means the algorithm takes a genuinely positive policy outcome and flips its sign before adding it to the final average. If Washington's policy created 5,000 jobs, a negative weight forces the model to count those jobs as losses.[3][6]

"Under heterogeneous treatment effects, the TWFE estimator can be negative even if the treatment effect is positive for every unit at every time period," de Chaisemartin and D'Haultfœuille warned in their analysis of the estimator's mechanics.[3]

When a policy's effect grows over time, subtracting an early adopter's trend accidentally subtracts the evolving treatment effect.

The researchers reviewed 116 highly cited papers published in top economics journals that relied on the standard algorithm. They found that 19 percent of the underlying comparisons in those papers were assigned negative weights by the mathematical engine.[3]

In a 2022 simulation published in the Journal of Financial Economics, researchers generated a fake dataset where a policy explicitly increased corporate outcomes by 18 percent. When they ran the standard algorithm, the negative weights overpowered the data, producing a final estimate of negative 5 percent.[5]

The severity of the distortion depends entirely on the variance in treatment timing. The more staggered the rollout, and the more the policy's effect changes over time, the more negative weights the algorithm generates.[1][5]

The Modern Solutions

The discovery of negative weights triggered an immediate methodological revolution in applied microeconomics. Researchers realized they could no longer trust the standard algorithm for any policy that rolled out in stages.[2][6]

Brantly Callaway and Pedro Sant'Anna developed a heterogeneity-robust estimator that completely bypasses the poisoned control groups. Their method manually constructs clean two-by-two comparisons, strictly forbidding the algorithm from ever using an already-treated state as a control.[2]

Liyang Sun and Sarah Abraham published a parallel solution for event studies in 2021. Their approach estimates a separate dynamic effect for each specific cohort of adopters, preventing the evolving effects from bleeding across the timeline.[4]

These modern estimators require researchers to explicitly define their counterfactuals rather than letting a black-box algorithm average them automatically. By forcing transparency, the new methods ensure that every mathematical weight remains strictly positive.[2][4][6]

In simulated datasets, negative weights can completely overpower the true data, flipping a positive 18 percent effect into a negative 5 percent estimate.

The transition to these robust estimators is now nearly complete across top academic journals. When an economist sits down to evaluate a staggered policy today, the two-way fixed effects engine remains firmly turned off.[5][6]

How we did this

Method
We decomposed the aggregate two-way fixed effects estimator into its constituent 2x2 difference-in-differences comparisons using the Goodman-Bacon theorem, isolating the specific variance-weighted term where an early-adopting cohort serves as the control for a late-adopting cohort.
What we found
When the early adopter's treatment effect grows dynamically by a magnitude larger than the late adopter's baseline change, the variance-weighting mechanism assigns a mathematical weight of less than zero to the late adopter's true effect, artificially shrinking the aggregate coefficient.
What we worked from
Limits of this analysis
This algebraic decomposition assumes a balanced panel with no unit-specific time trends, and the exact magnitude of the negative weight depends on the specific variance of treatment timing in a given dataset.

Key terms

Two-Way Fixed Effects (TWFE)
A statistical regression model that controls for constant differences between groups and uniform shocks across time to isolate the impact of a specific event.
Difference-in-Differences
An analytical method that estimates a policy's effect by comparing the change in outcomes for a treated group against the change in outcomes for an untreated control group.
Staggered Rollout
A scenario where different jurisdictions or groups adopt a policy or intervention at different points in time, rather than all at once.
Dynamic Treatment Effect
A policy impact that changes in magnitude over time, such as a tax credit that takes several years to reach its maximum economic benefit.

Frequently asked

Does this mean all older economic studies are wrong?

Not necessarily. Studies evaluating policies that rolled out all at once, or policies whose effects remained perfectly constant over time, are unaffected by the negative weighting issue.

How do the new estimators fix the problem?

Modern methods like Callaway and Sant'Anna explicitly block the algorithm from using already-treated units as controls, ensuring that only clean, uncontaminated counterfactuals are used in the calculation.

Can researchers still use two-way fixed effects?

Yes, but only if they can mathematically prove that their policy's effect does not change over time, a standard that is exceptionally difficult to meet in real-world economics.

Viewpoints in depth

Econometric Theorists

Focus on the mathematical proofs demonstrating that standard estimators fail under heterogeneous treatment effects.

Theorists approach the negative weighting problem as a fundamental failure of linear algebra. By decomposing the two-way fixed effects estimator into its constituent parts, they proved that the algorithm's reliance on variance-weighting inherently corrupts the estimate when treatment effects evolve over time. Their work shifted the discipline's focus from running regressions to explicitly defining the counterfactual comparisons that underpin causal inference.

Applied Microeconomists

Prioritize adopting new, robust estimators that provide accurate policy evaluations without discarding historical data.

For applied researchers, the discovery of negative weights threatened to invalidate decades of empirical findings on minimum wages, healthcare rollouts, and environmental regulations. Their immediate priority was implementing the new heterogeneity-robust estimators developed by Callaway, Sant'Anna, and others. Applied economists now routinely run both the flawed standard model and the corrected modern estimators side-by-side to demonstrate exactly how much the negative weights distorted their specific findings.

Policy Analysts

Rely on the corrected estimates to understand whether state-level interventions actually achieved their intended outcomes.

Policy analysts consume the final output of these models to advise lawmakers on future legislation. The revelation that a genuinely successful policy could be mathematically reported as a failure due to negative weights forced a reevaluation of several key policy debates. Analysts now demand that researchers explicitly account for staggered rollout dynamics before using academic findings to justify new government spending or regulatory changes.

Econometric Theorists 40%Applied Microeconomists 40%Policy Analysts 20%
Econometric Theorists
Focus on the mathematical proofs demonstrating that standard estimators fail under heterogeneous treatment effects.
Applied Microeconomists
Prioritize adopting new, robust estimators that provide accurate policy evaluations without discarding historical data.
Policy Analysts
Rely on the corrected estimates to understand whether state-level interventions actually achieved their intended outcomes.

Perspectives this story doesn't cover

  • Journal Editors
  • Historical Data Archivists

Sources

Source coverage

6 outlets

3 viewpoints surfaced

Econometric Theorists 40%Applied Microeconomists 40%Policy Analysts 20%
  1. [1]Journal of EconometricsEconometric Theorists

    Difference-in-differences with variation in treatment timing

    Read on Journal of Econometrics →
  2. [2]Journal of EconometricsEconometric Theorists

    Difference-in-Differences with multiple time periods

    Read on Journal of Econometrics →
  3. [3]American Economic ReviewEconometric Theorists

    Two-Way Fixed Effects Estimators with Heterogeneous Treatment Effects

    Read on American Economic Review →
  4. [4]Journal of EconometricsEconometric Theorists

    Estimating dynamic treatment effects in event studies with heterogeneous treatment effects

    Read on Journal of Econometrics →
  5. [5]Journal of Financial EconomicsApplied Microeconomists

    How much should we trust staggered difference-in-differences estimates?

    Read on Journal of Financial Economics →
  6. [6]Factlen Editorial TeamPolicy Analysts

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team →

Comments

Stay informed

Every angle. Every day.

Get Data & Analysis stories with full source coverage and perspective breakdowns, free every day.