Skip to main content
ExplainerCausal InferenceMethodology Explainer· 8 min read· in Data & Analysis

How Three Mathematical Conditions Validate an Instrumental Variable

When researchers cannot run a randomized trial, they rely on instrumental variables to isolate cause and effect. But the method's accuracy depends entirely on three rigid mathematical assumptions that actively amplify bias if broken.

By Sofia Matos

Econometric Purists 40%Applied Policy Evaluators 30%Epidemiological Methodologists 30%
Econometric Purists
Focuses on the strict mathematical boundaries of the Wald estimator and the Local Average Treatment Effect.
Applied Policy Evaluators
Prioritizes the practical application of quasi-experimental methods to measure the impact of real-world interventions.
Epidemiological Methodologists
Adapts instrumental variables to medical data to combat confounding by indication in observational health studies.

Perspectives this story doesn't cover

  • Machine Learning Theorists
  • Bayesian Statisticians
3
Core mathematical conditions required
4
Subgroups in the LATE compliance framework
2
Stages in the standard least squares regression
1920s
Decade the method was first deployed

Fast facts

  1. Instrumental variables isolate causal effects in observational data by leveraging naturally occurring shocks.
  2. The method requires the instrument to be strongly correlated with the treatment, known as the relevance condition.
  3. The exclusion restriction mandates that the instrument cannot affect the outcome through any secondary pathways.
  4. Monotonicity assumes the instrument pushes all affected individuals in the same direction, with no defiers.
  5. A weak instrument mathematically amplifies residual bias, often producing worse estimates than standard regressions.

How we got here

  1. 1920s

    Instrumental variables are first deployed to estimate supply and demand elasticities.

  2. 1996

    Economists formalize the Local Average Treatment Effect and the monotonicity assumption.

  3. 2001

    The National Bureau of Economic Research publishes a landmark framework for using instruments in natural experiments.

  4. 2023

    The World Bank updates its Development Impact Evaluation guidelines for quasi-experimental methods.

While a randomized controlled trial physically forces treatment assignment to be independent of unmeasured confounders, an instrumental variable attempts the exact same isolation using observational data and three rigid mathematical assumptions. When researchers cannot randomly assign a treatment—such as forcing individuals to smoke, drop out of high school, or adopt a specific medical device—they search for a naturally occurring shock that affects the treatment but has no direct link to the outcome. This shock, known as an instrument, allows statisticians to extract the causal effect from a tangle of correlated variables. But the validity of this extraction rests entirely on relevance, the exclusion restriction, and monotonicity. If any of the three fail, the instrument does not just fail to correct the bias; it actively amplifies it.[1][6]

The method of instrumental variables was first used in the 1920s to estimate supply and demand elasticities, later evolving to correct for measurement error in single-equation models. In a 2001 National Bureau of Economic Research working paper, economists Joshua Angrist and Alan Krueger detailed how the technique expanded into a primary tool for identifying causal relationships in natural experiments. The core problem the method solves is omitted variable bias. In standard observational data, a researcher attempting to measure the effect of education on earnings cannot observe a person's innate ability or family connections. Because those hidden variables affect both the likelihood of staying in school and the eventual salary, a standard ordinary least squares regression produces a distorted estimate.[6]

To bypass this distortion, researchers look for an instrument—a variable that shifts the probability of the treatment but is otherwise entirely disconnected from the outcome. Angrist and Krueger famously used the quarter of a student's birth as an instrument for educational attainment. Because compulsory schooling laws in the United States required students to remain in school until their 16th birthday, those born late in the year were forced to complete more schooling before they could legally drop out than those born early in the year. The quarter of birth acted as a random shock to education levels, completely independent of innate ability.[6]

The three mathematical conditions required to validate an instrumental variable.

The mechanics of this extraction rely on a mathematical operation often executed as Two-Stage Least Squares. In the first stage, the researcher regresses the treatment variable on the instrument to isolate the portion of the treatment that is driven solely by the external shock. In the second stage, the outcome is regressed on that isolated, predicted treatment value. However, as MIT economist Joshua Angrist and Princeton economist Alan Krueger explicitly warned in their 2001 National Bureau of Economic Research paper, "Instrumental variables estimates are not unbiased because they involve a ratio of random quantities, for which expectations need not exist or have a simple form." This ratio strips away the endogenous variation, but it requires massive sample sizes to converge on the true causal effect.[1][6]

This mathematical operation collapses if the first of the three core conditions—relevance—is not met. The relevance condition dictates that the instrument must have a strong, measurable effect on the treatment variable. A 2023 framework published by the World Bank's Development Impact Evaluation group emphasizes that finding a naturally occurring instrument is difficult, and finding one with a sufficiently strong correlation is even harder. If the instrument only shifts the probability of the treatment by 2 or 3 percentage points, it is classified as a weak instrument.[5]

The penalty for a weak instrument is severe. Because the instrumental variable estimator relies on a ratio, a weak correlation produces a denominator that approaches zero. A 2018 review in Emerging Themes in Epidemiology notes that a weak instrument exerts a multiplicative effect on any residual bias in the numerator. If an instrument shifts the treatment probability by just 5 percent, the denominator is 0.05. Any minor violation of the other assumptions is immediately multiplied by 20. Consequently, a weak instrument can produce an estimate that is significantly more biased than the flawed observational regression it was meant to replace.[2][5]

A weak instrument actively amplifies any minor violation of the exclusion restriction.

The second mandatory condition is the exclusion restriction. This assumption requires that the instrument affects the outcome exclusively through its effect on the treatment, with no alternative pathways. In the quarter-of-birth example, the exclusion restriction assumes that being born in the fourth quarter does not affect adult earnings through any mechanism other than the extra months of compulsory schooling. Unlike the relevance condition, which can be empirically tested by examining the first-stage regression, the exclusion restriction is fundamentally untestable. It relies entirely on the researcher's theoretical justification and domain knowledge.[1][4]

The second mandatory condition is the exclusion restriction.

When the exclusion restriction fails, the entire causal claim disintegrates. A 2018 analysis in Current Epidemiology Reports highlights that researchers often rely on falsification strategies to probe the plausibility of the exclusion restriction, even if they cannot prove it definitively. For example, researchers might test the instrument on a sub-population that is known to be immune to the treatment. If the instrument still produces a change in the outcome for that immune group, the exclusion restriction is violated, because the instrument must be operating through a back door.[3]

The third condition, monotonicity, was formalized in a landmark 1996 paper by Angrist, Guido Imbens, and Donald Rubin. Monotonicity requires that the instrument pushes all affected individuals in the same direction. In the language of causal inference, it assumes that there are no "defiers"—people who would take the treatment if not encouraged, but would refuse the treatment if encouraged. If an instrument is a financial subsidy for a medical procedure, monotonicity assumes that the subsidy might convince some people to get the procedure, but it will not cause anyone who was already planning to get the procedure to suddenly cancel it.[1][6]

When monotonicity holds, the instrumental variable isolates a specific parameter known as the Local Average Treatment Effect. The 1996 framework divided populations into four groups: always-takers, never-takers, compliers, and defiers. Because the instrument does not change the behavior of always-takers or never-takers, and because monotonicity assumes defiers do not exist, the resulting estimate applies exclusively to the compliers. It measures the causal effect only for the specific subset of the population whose behavior was changed by the instrument.[1][4]

The Local Average Treatment Effect measures the impact exclusively for the 'compliers'.

This local limitation is a frequent point of contention in policy evaluation. The World Bank notes that while an instrumental variable might accurately measure the impact of a credit scheme on the specific individuals who were induced to participate by a randomized encouragement, that estimate does not necessarily reflect the average treatment effect for the entire population. If the compliers differ systematically from the always-takers—perhaps because they are more risk-averse or have different baseline resources—extrapolating the local effect to a national policy rollout will yield inaccurate forecasts.[5]

Epidemiologists face similar constraints when adapting these econometric tools to medical research. A 2016 methodology review from Columbia University's Mailman School of Public Health outlines how instrumental variables are deployed to combat "confounding by indication." In observational medical data, patients who receive an aggressive treatment are often systematically sicker than those who do not, making the treatment appear harmful. Epidemiologists frequently use the prescribing preference of a specific physician or hospital as an instrument, assuming that a patient's assignment to a high-prescribing doctor is random and only affects their health through the increased likelihood of receiving the drug.[1][2]

In these medical applications, the terminology shifts—epidemiologists often refer to the "exchangeability" assumption rather than exogeneity—but the mathematical floor remains identical. The instrument must still be relevant (the doctor must actually prescribe the drug more often), it must satisfy the exclusion restriction (the doctor must not also provide better overall care that improves the outcome independently), and it must satisfy monotonicity (the doctor must not systematically deny the drug to patients who would have received it from a low-prescribing physician).[2][3]

Epidemiologists use instrumental variables to combat confounding by indication in observational medical data.

The strictness of these three conditions explains why valid instrumental variables are notoriously difficult to discover. The search requires finding a variable that is powerful enough to move human behavior but isolated enough to remain untainted by the complex web of socioeconomic or biological confounders. When researchers force an invalid instrument into a Two-Stage Least Squares model, the resulting output projects a false certainty, wrapping a biased estimate in the mathematical authority of causal inference.[4][6]

The frontier of instrumental variable research is now focused on quantifying the exact boundaries of these failures. Rather than treating the exclusion restriction and monotonicity as binary conditions that either hold perfectly or fail entirely, modern econometricians are developing bounds that calculate how much an estimate degrades when the assumptions are slightly violated. Until those sensitivity models become standard practice, the reliability of any instrumental variable estimate remains entirely dependent on the transparency of the researcher and the structural integrity of the naturally occurring shock.[3][6][7]

What we don’t know

  • The exact statistical threshold at which an instrument becomes too weak to reliably correct omitted variable bias in finite samples.
  • How to definitively prove the exclusion restriction, which remains a fundamentally untestable theoretical assumption.
  • The true average treatment effect for populations outside the specific 'complier' subgroup isolated by the instrument.

Sources

Source coverage

7 outlets

3 viewpoints surfaced

Econometric Purists 40%Applied Policy Evaluators 30%Epidemiological Methodologists 30%
  1. [1]Columbia UniversityEpidemiological Methodologists

    Instrumental Variables

    Read on Columbia University
  2. [2]Emerging Themes in EpidemiologyEpidemiological Methodologists

    An introduction to instrumental variable assumptions, validation and estimation

    Read on Emerging Themes in Epidemiology
  3. [3]Curr Epidemiol RepEpidemiological Methodologists

    Understanding the Assumptions Underlying Instrumental Variable Analyses: a Brief Review of Falsification Strategies and Related Tools

    Read on Curr Epidemiol Rep
  4. [4]IZA World of LaborEconometric Purists

    Using instrumental variables to establish causality

    Read on IZA World of Labor
  5. [5]The World BankApplied Policy Evaluators

    Instrumental Variables

    Read on The World Bank
  6. [6]National Bureau of Economic ResearchEconometric Purists

    Instrumental Variables and the Search for Identification: From Supply and Demand to Natural Experiments

    Read on National Bureau of Economic Research
  7. [7]Factlen Editorial TeamApplied Policy Evaluators

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team

Comments

Stay informed

Every angle. Every day.

Get Data & Analysis stories with full source coverage and perspective breakdowns delivered to your inbox.