Why Dropping High-VIF Controls to Fix Multicollinearity Induces Omitted Variable Bias
Discarding correlated variables artificially shrinks standard errors but mathematically forces the remaining coefficients to absorb the omitted effects. The resulting estimates become precisely and confidently incorrect.
By Ishani Patel
In short
- Multicollinearity inflates the variance of regression coefficients, driving up standard errors and rendering variables statistically insignificant.
- Despite the inflated variance, the Gauss-Markov theorem guarantees that the coefficient estimates themselves remain perfectly unbiased.
- Dropping a highly correlated control variable to fix the variance actively induces omitted variable bias, making the remaining coefficients systematically wrong.
In this article
When a data analyst runs a multiple regression model, they must decide exactly which variables to include and which to discard. They typically make this choice during the diagnostic phase, scanning for predictors that move together too closely. If two variables overlap, the analyst has the power to drop one before finalizing the model.
This overlap is known as multicollinearity, a condition where two or more independent variables are highly correlated. It is one of the most common diagnostic warnings in statistics. Analysts are routinely taught to fear it because it destabilizes the model's coefficient estimates, making them highly sensitive to small changes in the underlying data.
The standard tool for diagnosing this overlap is the Variance Inflation Factor, a metric developed by statistician Cuthbert Daniel. The VIF quantifies exactly how much the variance of a specific coefficient is inflated by its correlation with other predictors. It is calculated using the formula 1 / (1 - R²), where R² represents the correlation between the predictors.[2]
This auxiliary regression isolates the specific variable in question, testing how well it can be predicted by the remaining features in the dataset. If the other variables can predict it with near-perfect accuracy, the variable is providing almost no unique mathematical value to the overall model.
For example, if the auxiliary regression yields an R² of 0.90, the VIF calculation becomes 1 / (1 - 0.90), which equals 10. This means 90 percent of the variable's movement is already captured by the other variables. The model struggles to isolate the remaining 10 percent of unique information.
A VIF of 1 indicates zero correlation, meaning the variable's variance is completely uninflated. As the correlation approaches 100 percent, the VIF scales toward infinity. According to the Penn State Eberly College of Science, a VIF exceeding 4 warrants further investigation, while a value above 10 signals severe multicollinearity that requires correction.[4]
The Illusion of Broken Coefficients
When an analyst sees a VIF of 10, they know the variance of that coefficient is ten times larger than it would be if the variables were completely independent. This inflated variance produces massive standard errors. Consequently, the associated p-values spike, often rendering genuinely important variables statistically insignificant.
Faced with insignificant p-values, the instinctive reaction is to "fix" the model by dropping one of the correlated variables. If a dataset includes both a patient's weight and their body surface area—two metrics that correlate at roughly 0.87—an analyst might discard body surface area to bring the VIF back down to 1.[4]
However, this instinct fundamentally misunderstands what multicollinearity actually does to a model. While high correlation inflates the variance of the estimates, it does not bias the estimates themselves. The coefficients remain mathematically centered on the true population parameter, even if they are spread out over a wider distribution.
This guarantee is anchored in the Gauss-Markov theorem, a foundational principle of econometrics named after Carl Friedrich Gauss and Andrey Markov. "The Gauss-Markov theorem states that the ordinary least squares estimator is the best linear unbiased estimator," notes Wikipedia's summary of the theorem. As long as the model's errors are uncorrelated and have a mean of zero, the estimates remain unbiased.[3]
The Gauss-Markov theorem essentially promises that Ordinary Least Squares will not systematically overestimate or underestimate the true parameter. If an analyst were to draw thousands of different samples from the same population and run the same regression, the average of those highly variable coefficients would perfectly match the true population value.
This unbiasedness holds true regardless of how severely the variables are correlated, provided the correlation is not perfectly absolute. As long as there is even a fraction of independent variance between the predictors, the Ordinary Least Squares algorithm will eventually find the true causal center.
Because the estimates are unbiased, multicollinearity is essentially a problem of precision, not accuracy. The model is telling the analyst that it simply does not have enough independent information to confidently separate the effects of the two overlapping variables. It is an honest reflection of the data's limitations.
The model is essentially admitting that it cannot distinguish whether the first variable or the second variable is driving the outcome, because they always move together. The wide confidence intervals are a mathematical warning sign, urging the researcher not to draw overly confident conclusions from entangled data.
Inducing Omitted Variable Bias
When an analyst drops a high-VIF variable to artificially shrink the standard errors, they actively destroy that unbiasedness. If the discarded variable actually influences the dependent variable, removing it violates the core assumption of exogeneity. This triggers a mathematical consequence known as omitted variable bias.
"Omitted variable bias occurs when a statistical model fails to include one or more relevant variables," explains Kassiani Nikolopoulou in a 2022 methodology guide for Scribbr. By leaving out an important factor, the analyst forces the model to attribute the missing variable's effect to the variables that remain.[5]
The mechanics of this bias are strictly defined by the omitted variable bias formula. The bias added to the remaining coefficient equals the true effect of the omitted variable multiplied by the covariance between the included and omitted variables, divided by the variance of the included variable.[1]
This formula reveals that the magnitude of the bias is directly proportional to the strength of the correlation between the included and excluded variables. The tighter the two variables are linked, the more severely the remaining coefficient will be distorted when its partner is removed.
Because the two variables were highly correlated to begin with—which is why the VIF was high—their covariance is large. Therefore, dropping one mathematically guarantees that the remaining variable will absorb a massive share of the omitted variable's effect. The remaining coefficient becomes systematically, directionally wrong.
Consider a model estimating the effect of education on wages, which omits a highly correlated variable like innate ability. Because ability drives both education and wages, the education coefficient absorbs the hidden effect of ability. The model might estimate a $2.30 hourly wage increase per year of schooling, falsely inflating education's true causal impact.
If the true return to education is only $1.50 per hour, but the model estimates $2.30 because it absorbed the innate ability factor, the entire policy conclusion is compromised. The analyst has successfully lowered the VIF, but they have produced a fundamentally incorrect economic finding.
The direction of the omitted variable bias depends entirely on the signs of the underlying relationships. If the omitted variable has a positive effect on the outcome and a positive correlation with the included variable, the resulting bias will be strictly positive. The remaining coefficient will be artificially inflated.
Conversely, if the omitted variable has a negative effect on the outcome but a positive correlation with the included variable, the bias becomes negative. The included coefficient will be dragged downward, potentially flipping from a positive causal effect to a negative one, completely reversing the study's conclusions.
Trading Variance for Bias
By dropping a relevant control variable, the analyst has executed a disastrous statistical trade. They have successfully eliminated the inflated variance, resulting in a tight, highly significant standard error. But they have replaced it with a coefficient that is precisely and confidently incorrect.
In the context of causal inference, this trade-off is almost always a failure. An unbiased estimate with a wide confidence interval honestly communicates the uncertainty inherent in the data. A biased estimate with a narrow confidence interval actively misleads the researcher about the true relationship between the variables.
This distinction explains why econometricians and data scientists treat multicollinearity differently. In pure predictive machine learning, where the only goal is forecasting out-of-sample data, dropping collinear variables or using biased estimators can actually improve accuracy. The model does not care if a coefficient is causally wrong, as long as the prediction is right.
But in fields like economics, epidemiology, and policy analysis, the coefficient itself is the entire point of the study. If a health researcher wants to know the isolated effect of a specific drug, a biased coefficient could lead to dangerous medical guidelines. In these domains, unbiasedness cannot be sacrificed for a lower VIF.
Safer Alternatives to Dropping Variables
If dropping a variable induces omitted variable bias, analysts must rely on other methods to handle severe multicollinearity. The simplest, though often most expensive, solution is to collect more data. A larger sample size mathematically reduces the baseline variance, counteracting the inflation caused by the high VIF.
When collecting more data is impossible, researchers can combine the overlapping variables into a single composite index. If a survey includes highly correlated metrics like household income, occupational prestige, and years of education, they can be merged into a unified socioeconomic status score. This preserves the underlying information without triggering multicollinearity.
Another standard approach is Principal Component Analysis, a dimensionality reduction technique. PCA mathematically compresses the correlated variables into a set of new, entirely uncorrelated features. Because these new principal components are orthogonal to one another, they satisfy the regression assumptions perfectly and drop the VIF to exactly 1.
Because principal components are mathematically constructed to be independent, they completely bypass the multicollinearity problem. The analyst can regress the dependent variable on these new components, extracting the full predictive power of the original data without destabilizing the standard errors.
Finally, analysts can employ regularization techniques like Ridge regression. Ridge regression deliberately introduces a tiny, controlled amount of bias into the model by penalizing large coefficients. This mathematical constraint slashes the inflated variance and stabilizes the model, allowing it to digest correlated predictors without requiring manual deletion.
Finally, analysts can employ regularization techniques like Ridge regression.
By shrinking the coefficients toward zero, Ridge regression prevents any single variable from disproportionately dominating the model simply because it is correlated with another. The penalty term ensures that the model remains stable, even when the underlying data is highly collinear.
The most intellectually honest approach to multicollinearity is often to do nothing at all. If two variables are genuinely distinct concepts that both belong in the theoretical model, leaving them in preserves the model's structural integrity. The resulting wide standard errors are simply the price of statistical truth.
How we did this
- Method
- Mathematical derivation of the variance-bias trade-off when omitting a collinear regressor.
- What we found
- Dropping a relevant high-VIF variable mathematically guarantees that the remaining variable's coefficient will absorb the omitted effect, converting a purely variance-inflating issue into a systematic directional bias.
- What we worked from
- Limits of this analysis
- This applies to explanatory and causal inference models; in pure predictive machine learning, dropping collinear variables or using biased estimators like Ridge regression may improve out-of-sample accuracy.
Key terms
- Multicollinearity
- A statistical phenomenon where two or more independent variables in a regression model are highly correlated, making it difficult to isolate their individual effects.
- Variance Inflation Factor (VIF)
- A metric that quantifies how much the variance of an estimated regression coefficient is increased due to collinearity with other predictors.
- Omitted Variable Bias
- A systematic error that occurs when a statistical model leaves out a relevant variable, forcing the included variables to absorb its effect.
- Gauss-Markov Theorem
- A foundational mathematical proof stating that Ordinary Least Squares produces the best linear unbiased estimates, provided certain assumptions are met.
- Exogeneity
- The assumption that the independent variables in a regression model are not correlated with the model's error term.
- Ridge Regression
- A regularization technique that deliberately introduces a small amount of bias into a model to significantly reduce the variance of its coefficients.
Frequently asked
Does multicollinearity invalidate a regression model's predictions?
No. Multicollinearity inflates the standard errors of individual coefficients, making it hard to interpret their specific effects, but it does not degrade the model's overall predictive accuracy on the current dataset.
What is considered a dangerously high VIF score?
While thresholds vary by discipline, a VIF above 4 generally warrants investigation, and a VIF above 10 indicates severe multicollinearity that is heavily inflating the coefficient's variance.
Can omitted variable bias be fixed after a variable is dropped?
Not unless the dropped variable is reintroduced or an instrumental variable is used. Once a relevant variable is omitted, its effect is permanently absorbed into the remaining coefficients and the error term.
Why do machine learning models drop collinear variables if it causes bias?
Machine learning prioritizes overall predictive accuracy over causal truth. Introducing a small amount of bias to eliminate massive variance often results in better forecasts on new, unseen data.
Viewpoints in depth
Causal Inference Researchers
Economists and epidemiologists who prioritize unbiased coefficient estimates over tight confidence intervals.
For researchers attempting to isolate the true causal effect of a policy or medical intervention, unbiasedness is the paramount objective. They argue that a wide confidence interval honestly reflects the limitations of the dataset, whereas dropping a relevant variable to lower the VIF actively deceives the reader. They prefer to retain collinear controls and accept the inflated standard errors, ensuring the model remains structurally accurate.
Predictive Machine Learning Practitioners
Data scientists focused entirely on out-of-sample forecasting accuracy rather than causal explanation.
In pure predictive modeling, the individual coefficients do not need to reflect reality; they only need to combine to produce an accurate forecast. These practitioners frequently drop collinear variables or apply biased regularization techniques like Ridge and Lasso regression. By deliberately introducing a small amount of bias, they drastically reduce the model's variance, which reliably improves the model's performance on unseen test data.
Applied Statisticians
Pragmatic analysts who seek a middle ground through dimensionality reduction and composite indices.
Rather than choosing between inflated variance and omitted variable bias, applied statisticians often reshape the data itself. They advocate for techniques like Principal Component Analysis (PCA) or the creation of weighted indices. By mathematically compressing highly correlated variables into a single orthogonal feature, they eliminate the multicollinearity without discarding the underlying information, preserving both stability and theoretical integrity.
- Causal Inference Researchers
- Economists and epidemiologists who prioritize unbiased coefficient estimates over tight confidence intervals.
- Predictive Machine Learning Practitioners
- Data scientists focused entirely on out-of-sample forecasting accuracy rather than causal explanation.
- Applied Statisticians
- Pragmatic analysts who seek a middle ground through dimensionality reduction and composite indices.
Perspectives this story doesn't cover
- Software developers building automated feature-selection algorithms
Sources
[1]WikipediaOmitted-variable bias
Read on Wikipedia →
[2]WikipediaVariance inflation factor
Read on Wikipedia →
[3]WikipediaGauss–Markov theorem
Read on Wikipedia →
[4]Penn State Eberly College of ScienceApplied Statisticians10.4 - Multicollinearity
Read on Penn State Eberly College of Science →
[5]ScribbrOmitted Variable Bias | Definition, Examples & Causes
Read on Scribbr →
[6]Factlen Editorial TeamCausal Inference ResearchersSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
More in Data & Analysis
See all →Regression Math
The Five Assumptions That Guarantee Ordinary Least Squares is the Best Linear Unbiased Estimator
6 sources
Statistical Bias
How Immortal Time Bias Artificially Slashes Mortality Rates in Observational Medical Research
8 sources
Covariate Adjustment
Why Adjusting for Balanced Covariates Diverges Conditional and Marginal Odds Ratios
9 sources
Survival Analysis
How the Independent Censoring Assumption Artificially Inflates Absolute Risk in Medical Research
6 sources
Comments
Every angle. Every day.
Get Data & Analysis stories with full source coverage and perspective breakdowns, free every day.




