How Granger Causality Tests for Predictive Precedence, Not True Causal Influence
Introduced in 1969, the Granger causality test measures whether one time series helps forecast another, but its name often misleads analysts into assuming physical causation. The test strictly evaluates information flow and temporal precedence, leaving it vulnerable to confounding variables and lagged feedback loops.
- Forecasting Practitioners
- Value the test for its ability to identify leading indicators and reduce prediction error.
- Causal Inference Theorists
- Argue that true causation requires counterfactuals and interventions, not just temporal precedence.
- Statistical Methodologists
- Focus on the mathematical assumptions, such as stationarity and lag selection, required for the test to function.
Perspectives this story doesn't cover
- Policy Makers
- Epidemiologists
What we don’t know
- Whether a statistically significant Granger causality result in a specific dataset is driven by a true physical mechanism or an unobserved confounding variable.
- The exact lag length that perfectly captures the temporal relationship between two variables without overfitting the model.
- How to fully adapt the linear Granger causality framework to capture complex, non-linear relationships in highly chaotic systems.
In 1969, within the pages of the journal Econometrica, British econometrician Clive Granger proposed a mathematical workaround for one of the oldest philosophical traps in statistics. The trap was the impossibility of proving true causation from observational data. Granger's solution was to shift the goalpost. Instead of asking whether variable X physically brings about variable Y, he asked a narrower, highly specific question: does knowing the history of X improve our ability to forecast the future of Y?[1]
This concept, which Granger initially termed 'temporally related' but which the wider scientific community quickly branded 'Granger causality,' became a cornerstone of time series analysis. It provided a computationally simple way to measure information flow between two sequences of data points ordered chronologically. However, the adoption of the word 'causality' created a linguistic trap that continues to snare data scientists, economists, and medical researchers decades later.[5][6]
To understand why Granger causality is not true causality, one must look at the mechanics of the test itself. The procedure relies on Vector Autoregression (VAR). An analyst first builds a restricted model, attempting to predict the current value of variable Y using only its own past values, testing anywhere from 1 to 10 previous time steps, known as lags. They then build an unrestricted model, which includes both the past values of Y and the past values of variable X.[3][4]
The mathematical verdict relies on an F-test, which compares the predictive accuracy of the two models. If the unrestricted model—the one containing X's history—produces a statistically significant reduction in forecasting error, typically measured against a standard alpha threshold of 0.05, the test rejects the null hypothesis. The analyst concludes that X 'Granger-causes' Y. The past values of X contain unique information about Y's future that Y's own history does not already possess.[4][6]
This is a statement about predictive precedence, not physical mechanism. True causal inference, as defined in modern econometrics and epidemiology, requires counterfactuals: proving that an intervention changing X would physically alter the trajectory of Y. The Granger test requires no such intervention. It merely observes the sequence of events, making it highly vulnerable to confounding variables.[2][5]
This is a statement about predictive precedence, not physical mechanism.
The most common failure mode occurs when a third, unobserved variable drives both X and Y, but at different speeds. Imagine a scenario where a central bank secretly decides to raise interest rates. Bond markets (variable X) might react within 50 milliseconds, while consumer mortgage rates (variable Y) take 14 to 21 days to adjust. A Granger causality test would confidently declare that bond market movements 'cause' mortgage rate hikes. In reality, both are reacting to the central bank's hidden decision; intervening to artificially manipulate the bond market would not force mortgage rates to rise.[5][6]
A second major limitation is lagged reverse causality. The test checks whether X's lags predict Y, but it cannot rule out that the true causal arrow runs from Y to X with a delay longer than the analyst has modeled. If Y physically causes X, but takes 12 months to do so, and the analyst only tests a 3-month lag, the test might register X as predicting Y simply because X is reacting to an earlier, unmeasured value of Y that has not yet fully worked through the system.[3][4]
Furthermore, the mathematical validity of the test rests on strict assumptions about the underlying data. Both time series must be stationary, meaning their statistical properties—like mean and variance—do not change over time. If the data contains trends or seasonal cycles, the test can produce spurious correlations, falsely identifying predictive precedence where none exists. Analysts must often difference the data—subtracting the value at time T-1 from the value at time T—to achieve stationarity before running the test, a step that can obscure long-term relationships.[3][4]
Despite these limitations, the test remains a powerful tool when used for its original, intended purpose: feature selection in forecasting. In fields ranging from macroeconomic nowcasting to machine learning, knowing which variables contain early signals of a target outcome is immensely valuable, regardless of the underlying physical mechanism.[5][6]
The danger arises only when the output of a forecasting filter is presented as an epidemiological or policy proof. 'Of course, many ridiculous papers appeared,' Granger noted in his 2003 Nobel lecture, referring to studies that assumed his statistical filter proved physical mechanisms. The mathematical reality remains unchanged: temporal precedence is a necessary condition for true causality, but it is never a sufficient one.[2][5]
Key points
- Granger causality tests whether the past values of one variable improve the forecast of another, not whether one physically causes the other.
- The test relies on Vector Autoregression (VAR) to compare a restricted predictive model against an unrestricted one.
- Unobserved confounding variables can easily trick the test if they influence two variables at different speeds.
- The underlying time series data must be stationary for the mathematical results to be valid.
- Clive Granger himself warned against interpreting the test as proof of a physical mechanism, noting it only measures temporal precedence.
How we got here
1969
Clive Granger publishes his seminal paper in Econometrica, introducing the concept of testing for predictive precedence using lagged variables.
1977
Granger attempts to clarify the terminology, suggesting the relationship is better described as 'temporally related' rather than causal.
2003
Granger is awarded the Nobel Memorial Prize in Economic Sciences, using his lecture to warn against 'ridiculous' applications of the test.
2011
The National Bureau of Economic Research publishes comprehensive reviews contrasting Granger's predictive test with true counterfactual causal inference.
Sources
[1]EconometricaForecasting PractitionersInvestigating Causal Relations by Econometric Models and Cross-Spectral Methods
Read on Econometrica →
[2]National Bureau of Economic ResearchCausal Inference TheoristsEconomics, History, and Causation
Read on National Bureau of Economic Research →
[3]ScholarpediaStatistical MethodologistsGranger causality
Read on Scholarpedia →
[4]National Institutes of HealthStatistical MethodologistsGranger Causality: A Review and Recent Advances
Read on National Institutes of Health →
[5]WikipediaStatistical MethodologistsGranger causality
Read on Wikipedia →
[6]Factlen Editorial TeamForecasting PractitionersSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
Comments
More in Data & Analysis
See all →Experiment Design
How the Minimum Detectable Effect, Statistical Power, and Alpha Determine the Required Sample Size for an A/B Test
6 sources
Imbalanced Data
Evidence Pack: The Accuracy and Trade-Offs of SMOTE Versus Class Weights in Imbalanced Data
6 sources
Ensemble Methods
Evidence Pack: How Bagging Reduces Variance While Boosting Reduces Bias in Ensemble Models
7 sources
Parkinson's Research
Machine Learning Model Predicts Parkinson's Disease Up to Seven Years Before Symptom Onset Via Blood Biomarkers
2 sources
Every angle. Every day.
Get Data & Analysis stories with full source coverage and perspective breakdowns delivered to your inbox.




