Already-Treated Controls Subtract Evolving Treatment Effects: Why Two-Way Fixed Effects Assign Negative Weights in Staggered Rollouts
When policies roll out at different times, standard statistical models mistakenly use early adopters as control groups for late adopters. If the policy's impact grows over time, this mathematical quirk subtracts the evolving effect, assigning negative weights that can artificially reverse the policy's true outcome.
By Sofia Matos
In short
- Standard statistical models mistakenly use early policy adopters as control groups for late adopters during staggered rollouts.
- If a policy's effect grows over time, subtracting the early adopter's trend mathematically assigns a negative weight to the true outcome.
- Modern heterogeneity-robust estimators solve this by strictly forbidding the algorithm from using already-treated units as counterfactuals.
When an economist sits down to evaluate a policy that rolls out state by state, they must choose a mathematical engine to isolate the effect. The researcher decides which variables to hold constant, what timeline to measure, and how to separate the policy's impact from background noise.[6]
For twenty years, the default engine for this task was the two-way fixed effects regression. This statistical model became the undisputed standard in applied microeconomics, appearing in thousands of papers evaluating everything from minimum wage hikes to healthcare expansions.[1][3]
The model relies on a foundational logic called difference-in-differences. If California raises its minimum wage in 2018 and Nevada does not, the researcher compares the change in California's employment to the change in Nevada's employment over the exact same period.[6]
Nevada serves as the counterfactual, representing what would have happened in California had the policy never passed. By subtracting Nevada's baseline trend from California's observed outcome, the model isolates the specific effect of the wage hike.[1]
The Staggered Rollout Trap
This simple two-by-two comparison works flawlessly when a policy happens exactly once. But real-world legislation rarely follows such a clean schedule, and policies typically cascade across jurisdictions over several years.[6]
When states adopt a policy at different times, researchers face a staggered rollout. California might act in 2018, Oregon in 2019, and Washington in 2020, creating a complex matrix of treatment timing that a simple two-by-two grid cannot accommodate.[1][2]
To handle this complexity, researchers feed the entire dataset into a two-way fixed effects regression. The algorithm mathematically averages all the available comparisons into a single, neat coefficient that represents the policy's overall average effect.[3]
The model accomplishes this by controlling for fixed differences between states and fixed shocks across time. It assumes that once these baseline differences are stripped away, every remaining variation in the data represents the pure impact of the policy.[1][3]
But in 2021, econometricians discovered a fatal flaw in how the algorithm constructs that average. The model does not just compare treated states to states that never adopted the policy; it also compares newly treated states to states that adopted the policy years earlier.[1][4]
Unpacking the Black Box
Andrew Goodman-Bacon published a landmark theorem in the Journal of Econometrics that cracked open the algorithm's black box. He proved that the single coefficient is actually a weighted average of every possible two-by-two comparison in the dataset.[1]
"The TWFE estimator is a weighted average of all two-group/two-period DiD estimators in the data," Goodman-Bacon wrote. His decomposition revealed exactly which states the algorithm was pairing together behind the scenes.[1]
Some of these pairings make intuitive sense, like comparing a state treated in 2019 against a state that is never treated. The algorithm assigns heavy mathematical weights to these clean comparisons, anchoring the final estimate in solid counterfactuals.[1][6]
Other pairings compare early adopters to late adopters before the late adopters pass the policy. Because neither state has changed its status during that specific window, this comparison also yields a valid measure of the policy's impact.[1]
The critical failure occurs in the final type of pairing. The algorithm routinely uses early adopters as the control group for late adopters, comparing Washington's 2020 policy change against California's 2020 baseline.[1][2]
The Poisoned Control Group
Using California as a control group in 2020 requires a massive mathematical assumption. The algorithm must assume that California's 2018 policy has already reached its maximum effect and that its current trajectory perfectly mirrors a world without the policy.[3][4]
In reality, economic policies almost always have dynamic treatment effects that evolve over time. A new tax credit might take three years to reach peak enrollment, meaning the policy's impact grows steadily larger with each passing year.[4][6]
If California's policy effect is still growing in 2020, its baseline trend is fundamentally poisoned. The state is experiencing a delayed surge from its own 2018 legislation, making it a terrible proxy for what Washington would have looked like without the 2020 law.[1][4]
When the algorithm subtracts California's poisoned trend from Washington's outcome, it accidentally subtracts the evolving treatment effect. The model interprets California's delayed surge as a background economic shock, and penalizes Washington's estimate accordingly.[1][3]
Clément de Chaisemartin and Xavier D'Haultfœuille quantified this disaster in the American Economic Review in 2020. They demonstrated that this subtraction mechanism can mathematically assign a negative weight to the true treatment effect of the late-adopting state.[3]
The Mathematics of Negative Weights
A negative weight means the algorithm takes a genuinely positive policy outcome and flips its sign before adding it to the final average. If Washington's policy created 5,000 jobs, a negative weight forces the model to count those jobs as losses.[3][6]
"Under heterogeneous treatment effects, the TWFE estimator can be negative even if the treatment effect is positive for every unit at every time period," de Chaisemartin and D'Haultfœuille warned in their analysis of the estimator's mechanics.[3]
The researchers reviewed 116 highly cited papers published in top economics journals that relied on the standard algorithm. They found that 19 percent of the underlying comparisons in those papers were assigned negative weights by the mathematical engine.[3]
In a 2022 simulation published in the Journal of Financial Economics, researchers generated a fake dataset where a policy explicitly increased corporate outcomes by 18 percent. When they ran the standard algorithm, the negative weights overpowered the data, producing a final estimate of negative 5 percent.[5]
The severity of the distortion depends entirely on the variance in treatment timing. The more staggered the rollout, and the more the policy's effect changes over time, the more negative weights the algorithm generates.[1][5]
The Modern Solutions
The discovery of negative weights triggered an immediate methodological revolution in applied microeconomics. Researchers realized they could no longer trust the standard algorithm for any policy that rolled out in stages.[2][6]
Brantly Callaway and Pedro Sant'Anna developed a heterogeneity-robust estimator that completely bypasses the poisoned control groups. Their method manually constructs clean two-by-two comparisons, strictly forbidding the algorithm from ever using an already-treated state as a control.[2]
Liyang Sun and Sarah Abraham published a parallel solution for event studies in 2021. Their approach estimates a separate dynamic effect for each specific cohort of adopters, preventing the evolving effects from bleeding across the timeline.[4]
These modern estimators require researchers to explicitly define their counterfactuals rather than letting a black-box algorithm average them automatically. By forcing transparency, the new methods ensure that every mathematical weight remains strictly positive.[2][4][6]
How we did this
- Method
- We decomposed the aggregate two-way fixed effects estimator into its constituent 2x2 difference-in-differences comparisons using the Goodman-Bacon theorem, isolating the specific variance-weighted term where an early-adopting cohort serves as the control for a late-adopting cohort.
- What we found
- When the early adopter's treatment effect grows dynamically by a magnitude larger than the late adopter's baseline change, the variance-weighting mechanism assigns a mathematical weight of less than zero to the late adopter's true effect, artificially shrinking the aggregate coefficient.
- What we worked from
- Early-adopter post-treatment trend component: ΔY_early — Journal of Econometrics
- Late-adopter treatment effect: ΔY_late — Journal of Econometrics
- Limits of this analysis
- This algebraic decomposition assumes a balanced panel with no unit-specific time trends, and the exact magnitude of the negative weight depends on the specific variance of treatment timing in a given dataset.
Key terms
- Two-Way Fixed Effects (TWFE)
- A statistical regression model that controls for constant differences between groups and uniform shocks across time to isolate the impact of a specific event.
- Difference-in-Differences
- An analytical method that estimates a policy's effect by comparing the change in outcomes for a treated group against the change in outcomes for an untreated control group.
- Staggered Rollout
- A scenario where different jurisdictions or groups adopt a policy or intervention at different points in time, rather than all at once.
- Dynamic Treatment Effect
- A policy impact that changes in magnitude over time, such as a tax credit that takes several years to reach its maximum economic benefit.
Frequently asked
Does this mean all older economic studies are wrong?
Not necessarily. Studies evaluating policies that rolled out all at once, or policies whose effects remained perfectly constant over time, are unaffected by the negative weighting issue.
How do the new estimators fix the problem?
Modern methods like Callaway and Sant'Anna explicitly block the algorithm from using already-treated units as controls, ensuring that only clean, uncontaminated counterfactuals are used in the calculation.
Can researchers still use two-way fixed effects?
Yes, but only if they can mathematically prove that their policy's effect does not change over time, a standard that is exceptionally difficult to meet in real-world economics.
Viewpoints in depth
Econometric Theorists
Focus on the mathematical proofs demonstrating that standard estimators fail under heterogeneous treatment effects.
Theorists approach the negative weighting problem as a fundamental failure of linear algebra. By decomposing the two-way fixed effects estimator into its constituent parts, they proved that the algorithm's reliance on variance-weighting inherently corrupts the estimate when treatment effects evolve over time. Their work shifted the discipline's focus from running regressions to explicitly defining the counterfactual comparisons that underpin causal inference.
Applied Microeconomists
Prioritize adopting new, robust estimators that provide accurate policy evaluations without discarding historical data.
For applied researchers, the discovery of negative weights threatened to invalidate decades of empirical findings on minimum wages, healthcare rollouts, and environmental regulations. Their immediate priority was implementing the new heterogeneity-robust estimators developed by Callaway, Sant'Anna, and others. Applied economists now routinely run both the flawed standard model and the corrected modern estimators side-by-side to demonstrate exactly how much the negative weights distorted their specific findings.
Policy Analysts
Rely on the corrected estimates to understand whether state-level interventions actually achieved their intended outcomes.
Policy analysts consume the final output of these models to advise lawmakers on future legislation. The revelation that a genuinely successful policy could be mathematically reported as a failure due to negative weights forced a reevaluation of several key policy debates. Analysts now demand that researchers explicitly account for staggered rollout dynamics before using academic findings to justify new government spending or regulatory changes.
- Econometric Theorists
- Focus on the mathematical proofs demonstrating that standard estimators fail under heterogeneous treatment effects.
- Applied Microeconomists
- Prioritize adopting new, robust estimators that provide accurate policy evaluations without discarding historical data.
- Policy Analysts
- Rely on the corrected estimates to understand whether state-level interventions actually achieved their intended outcomes.
Perspectives this story doesn't cover
- Journal Editors
- Historical Data Archivists
Sources
[1]Journal of EconometricsEconometric TheoristsDifference-in-differences with variation in treatment timing
Read on Journal of Econometrics →
[2]Journal of EconometricsEconometric TheoristsDifference-in-Differences with multiple time periods
Read on Journal of Econometrics →
[3]American Economic ReviewEconometric TheoristsTwo-Way Fixed Effects Estimators with Heterogeneous Treatment Effects
Read on American Economic Review →
[4]Journal of EconometricsEconometric TheoristsEstimating dynamic treatment effects in event studies with heterogeneous treatment effects
Read on Journal of Econometrics →
[5]Journal of Financial EconomicsApplied MicroeconomistsHow much should we trust staggered difference-in-differences estimates?
Read on Journal of Financial Economics →
[6]Factlen Editorial TeamPolicy AnalystsSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
More in Data & Analysis
See all →Statistical Bias
How Immortal Time Bias Artificially Slashes Mortality Rates in Observational Medical Research
8 sources
Covariate Adjustment
Why Adjusting for Balanced Covariates Diverges Conditional and Marginal Odds Ratios
9 sources
Survival Analysis
How the Independent Censoring Assumption Artificially Inflates Absolute Risk in Medical Research
6 sources
Ecosystem Restoration
Neglecting Degraded Land Costs Up to $6.3 Trillion Yearly, UN Forest Assessment Warns
5 sources
Comments
Every angle. Every day.
Get Data & Analysis stories with full source coverage and perspective breakdowns, free every day.




