Skip to main content
ExplainerRegression DiagnosticsEconometrics· 9 min read· in Data & Analysis

Why Dropping High-VIF Controls to Fix Multicollinearity Induces Omitted Variable Bias

Discarding correlated variables artificially shrinks standard errors but mathematically forces the remaining coefficients to absorb the omitted effects. The resulting estimates become precisely and confidently incorrect.

By Ishani Patel

In short

  • Multicollinearity inflates the variance of regression coefficients, driving up standard errors and rendering variables statistically insignificant.
  • Despite the inflated variance, the Gauss-Markov theorem guarantees that the coefficient estimates themselves remain perfectly unbiased.
  • Dropping a highly correlated control variable to fix the variance actively induces omitted variable bias, making the remaining coefficients systematically wrong.

When a data analyst runs a multiple regression model, they must decide exactly which variables to include and which to discard. They typically make this choice during the diagnostic phase, scanning for predictors that move together too closely. If two variables overlap, the analyst has the power to drop one before finalizing the model.

This overlap is known as multicollinearity, a condition where two or more independent variables are highly correlated. It is one of the most common diagnostic warnings in statistics. Analysts are routinely taught to fear it because it destabilizes the model's coefficient estimates, making them highly sensitive to small changes in the underlying data.

The standard tool for diagnosing this overlap is the Variance Inflation Factor, a metric developed by statistician Cuthbert Daniel. The VIF quantifies exactly how much the variance of a specific coefficient is inflated by its correlation with other predictors. It is calculated using the formula 1 / (1 - R²), where R² represents the correlation between the predictors.[2]

This auxiliary regression isolates the specific variable in question, testing how well it can be predicted by the remaining features in the dataset. If the other variables can predict it with near-perfect accuracy, the variable is providing almost no unique mathematical value to the overall model.

The Variance Inflation Factor quantifies exactly how much a coefficient's variance is inflated by correlation with other predictors.

For example, if the auxiliary regression yields an R² of 0.90, the VIF calculation becomes 1 / (1 - 0.90), which equals 10. This means 90 percent of the variable's movement is already captured by the other variables. The model struggles to isolate the remaining 10 percent of unique information.

A VIF of 1 indicates zero correlation, meaning the variable's variance is completely uninflated. As the correlation approaches 100 percent, the VIF scales toward infinity. According to the Penn State Eberly College of Science, a VIF exceeding 4 warrants further investigation, while a value above 10 signals severe multicollinearity that requires correction.[4]

The Illusion of Broken Coefficients

When an analyst sees a VIF of 10, they know the variance of that coefficient is ten times larger than it would be if the variables were completely independent. This inflated variance produces massive standard errors. Consequently, the associated p-values spike, often rendering genuinely important variables statistically insignificant.

Faced with insignificant p-values, the instinctive reaction is to "fix" the model by dropping one of the correlated variables. If a dataset includes both a patient's weight and their body surface area—two metrics that correlate at roughly 0.87—an analyst might discard body surface area to bring the VIF back down to 1.[4]

However, this instinct fundamentally misunderstands what multicollinearity actually does to a model. While high correlation inflates the variance of the estimates, it does not bias the estimates themselves. The coefficients remain mathematically centered on the true population parameter, even if they are spread out over a wider distribution.

This guarantee is anchored in the Gauss-Markov theorem, a foundational principle of econometrics named after Carl Friedrich Gauss and Andrey Markov. "The Gauss-Markov theorem states that the ordinary least squares estimator is the best linear unbiased estimator," notes Wikipedia's summary of the theorem. As long as the model's errors are uncorrelated and have a mean of zero, the estimates remain unbiased.[3]

When a relevant variable is omitted, its causal effect is absorbed by the remaining correlated variables, systematically biasing their coefficients.

The Gauss-Markov theorem essentially promises that Ordinary Least Squares will not systematically overestimate or underestimate the true parameter. If an analyst were to draw thousands of different samples from the same population and run the same regression, the average of those highly variable coefficients would perfectly match the true population value.

This unbiasedness holds true regardless of how severely the variables are correlated, provided the correlation is not perfectly absolute. As long as there is even a fraction of independent variance between the predictors, the Ordinary Least Squares algorithm will eventually find the true causal center.

Because the estimates are unbiased, multicollinearity is essentially a problem of precision, not accuracy. The model is telling the analyst that it simply does not have enough independent information to confidently separate the effects of the two overlapping variables. It is an honest reflection of the data's limitations.

The model is essentially admitting that it cannot distinguish whether the first variable or the second variable is driving the outcome, because they always move together. The wide confidence intervals are a mathematical warning sign, urging the researcher not to draw overly confident conclusions from entangled data.

Inducing Omitted Variable Bias

When an analyst drops a high-VIF variable to artificially shrink the standard errors, they actively destroy that unbiasedness. If the discarded variable actually influences the dependent variable, removing it violates the core assumption of exogeneity. This triggers a mathematical consequence known as omitted variable bias.

"Omitted variable bias occurs when a statistical model fails to include one or more relevant variables," explains Kassiani Nikolopoulou in a 2022 methodology guide for Scribbr. By leaving out an important factor, the analyst forces the model to attribute the missing variable's effect to the variables that remain.[5]

Dropping a variable trades an unbiased estimate with high variance for a biased estimate with low variance—a dangerous trade in causal inference.

The mechanics of this bias are strictly defined by the omitted variable bias formula. The bias added to the remaining coefficient equals the true effect of the omitted variable multiplied by the covariance between the included and omitted variables, divided by the variance of the included variable.[1]

This formula reveals that the magnitude of the bias is directly proportional to the strength of the correlation between the included and excluded variables. The tighter the two variables are linked, the more severely the remaining coefficient will be distorted when its partner is removed.

Because the two variables were highly correlated to begin with—which is why the VIF was high—their covariance is large. Therefore, dropping one mathematically guarantees that the remaining variable will absorb a massive share of the omitted variable's effect. The remaining coefficient becomes systematically, directionally wrong.

Consider a model estimating the effect of education on wages, which omits a highly correlated variable like innate ability. Because ability drives both education and wages, the education coefficient absorbs the hidden effect of ability. The model might estimate a $2.30 hourly wage increase per year of schooling, falsely inflating education's true causal impact.

If the true return to education is only $1.50 per hour, but the model estimates $2.30 because it absorbed the innate ability factor, the entire policy conclusion is compromised. The analyst has successfully lowered the VIF, but they have produced a fundamentally incorrect economic finding.

The direction of the omitted variable bias depends entirely on the signs of the underlying relationships. If the omitted variable has a positive effect on the outcome and a positive correlation with the included variable, the resulting bias will be strictly positive. The remaining coefficient will be artificially inflated.

Statisticians rely on dimensionality reduction and regularization to handle multicollinearity without resorting to variable deletion.

Conversely, if the omitted variable has a negative effect on the outcome but a positive correlation with the included variable, the bias becomes negative. The included coefficient will be dragged downward, potentially flipping from a positive causal effect to a negative one, completely reversing the study's conclusions.

Trading Variance for Bias

By dropping a relevant control variable, the analyst has executed a disastrous statistical trade. They have successfully eliminated the inflated variance, resulting in a tight, highly significant standard error. But they have replaced it with a coefficient that is precisely and confidently incorrect.

In the context of causal inference, this trade-off is almost always a failure. An unbiased estimate with a wide confidence interval honestly communicates the uncertainty inherent in the data. A biased estimate with a narrow confidence interval actively misleads the researcher about the true relationship between the variables.

This distinction explains why econometricians and data scientists treat multicollinearity differently. In pure predictive machine learning, where the only goal is forecasting out-of-sample data, dropping collinear variables or using biased estimators can actually improve accuracy. The model does not care if a coefficient is causally wrong, as long as the prediction is right.

But in fields like economics, epidemiology, and policy analysis, the coefficient itself is the entire point of the study. If a health researcher wants to know the isolated effect of a specific drug, a biased coefficient could lead to dangerous medical guidelines. In these domains, unbiasedness cannot be sacrificed for a lower VIF.

Safer Alternatives to Dropping Variables

If dropping a variable induces omitted variable bias, analysts must rely on other methods to handle severe multicollinearity. The simplest, though often most expensive, solution is to collect more data. A larger sample size mathematically reduces the baseline variance, counteracting the inflation caused by the high VIF.

Illustration: The Gauss-Markov theorem guarantees that Ordinary Least Squares estimates remain unbiased even under severe multicollinearity.

When collecting more data is impossible, researchers can combine the overlapping variables into a single composite index. If a survey includes highly correlated metrics like household income, occupational prestige, and years of education, they can be merged into a unified socioeconomic status score. This preserves the underlying information without triggering multicollinearity.

Another standard approach is Principal Component Analysis, a dimensionality reduction technique. PCA mathematically compresses the correlated variables into a set of new, entirely uncorrelated features. Because these new principal components are orthogonal to one another, they satisfy the regression assumptions perfectly and drop the VIF to exactly 1.

Because principal components are mathematically constructed to be independent, they completely bypass the multicollinearity problem. The analyst can regress the dependent variable on these new components, extracting the full predictive power of the original data without destabilizing the standard errors.

Finally, analysts can employ regularization techniques like Ridge regression. Ridge regression deliberately introduces a tiny, controlled amount of bias into the model by penalizing large coefficients. This mathematical constraint slashes the inflated variance and stabilizes the model, allowing it to digest correlated predictors without requiring manual deletion.

Finally, analysts can employ regularization techniques like Ridge regression.

By shrinking the coefficients toward zero, Ridge regression prevents any single variable from disproportionately dominating the model simply because it is correlated with another. The penalty term ensures that the model remains stable, even when the underlying data is highly collinear.

The most intellectually honest approach to multicollinearity is often to do nothing at all. If two variables are genuinely distinct concepts that both belong in the theoretical model, leaving them in preserves the model's structural integrity. The resulting wide standard errors are simply the price of statistical truth.

How we did this

Method
Mathematical derivation of the variance-bias trade-off when omitting a collinear regressor.
What we found
Dropping a relevant high-VIF variable mathematically guarantees that the remaining variable's coefficient will absorb the omitted effect, converting a purely variance-inflating issue into a systematic directional bias.
What we worked from
  • Variance Inflation Factor formula: VIF = 1 / (1 - R²) — Wikipedia
  • Omitted Variable Bias definition: Bias occurs when a relevant variable is excluded, forcing included variables to absorb its effect. — Scribbr
Limits of this analysis
This applies to explanatory and causal inference models; in pure predictive machine learning, dropping collinear variables or using biased estimators like Ridge regression may improve out-of-sample accuracy.

Key terms

Multicollinearity
A statistical phenomenon where two or more independent variables in a regression model are highly correlated, making it difficult to isolate their individual effects.
Variance Inflation Factor (VIF)
A metric that quantifies how much the variance of an estimated regression coefficient is increased due to collinearity with other predictors.
Omitted Variable Bias
A systematic error that occurs when a statistical model leaves out a relevant variable, forcing the included variables to absorb its effect.
Gauss-Markov Theorem
A foundational mathematical proof stating that Ordinary Least Squares produces the best linear unbiased estimates, provided certain assumptions are met.
Exogeneity
The assumption that the independent variables in a regression model are not correlated with the model's error term.
Ridge Regression
A regularization technique that deliberately introduces a small amount of bias into a model to significantly reduce the variance of its coefficients.

Frequently asked

Does multicollinearity invalidate a regression model's predictions?

No. Multicollinearity inflates the standard errors of individual coefficients, making it hard to interpret their specific effects, but it does not degrade the model's overall predictive accuracy on the current dataset.

What is considered a dangerously high VIF score?

While thresholds vary by discipline, a VIF above 4 generally warrants investigation, and a VIF above 10 indicates severe multicollinearity that is heavily inflating the coefficient's variance.

Can omitted variable bias be fixed after a variable is dropped?

Not unless the dropped variable is reintroduced or an instrumental variable is used. Once a relevant variable is omitted, its effect is permanently absorbed into the remaining coefficients and the error term.

Why do machine learning models drop collinear variables if it causes bias?

Machine learning prioritizes overall predictive accuracy over causal truth. Introducing a small amount of bias to eliminate massive variance often results in better forecasts on new, unseen data.

Viewpoints in depth

Causal Inference Researchers

Economists and epidemiologists who prioritize unbiased coefficient estimates over tight confidence intervals.

For researchers attempting to isolate the true causal effect of a policy or medical intervention, unbiasedness is the paramount objective. They argue that a wide confidence interval honestly reflects the limitations of the dataset, whereas dropping a relevant variable to lower the VIF actively deceives the reader. They prefer to retain collinear controls and accept the inflated standard errors, ensuring the model remains structurally accurate.

Predictive Machine Learning Practitioners

Data scientists focused entirely on out-of-sample forecasting accuracy rather than causal explanation.

In pure predictive modeling, the individual coefficients do not need to reflect reality; they only need to combine to produce an accurate forecast. These practitioners frequently drop collinear variables or apply biased regularization techniques like Ridge and Lasso regression. By deliberately introducing a small amount of bias, they drastically reduce the model's variance, which reliably improves the model's performance on unseen test data.

Applied Statisticians

Pragmatic analysts who seek a middle ground through dimensionality reduction and composite indices.

Rather than choosing between inflated variance and omitted variable bias, applied statisticians often reshape the data itself. They advocate for techniques like Principal Component Analysis (PCA) or the creation of weighted indices. By mathematically compressing highly correlated variables into a single orthogonal feature, they eliminate the multicollinearity without discarding the underlying information, preserving both stability and theoretical integrity.

Causal Inference Researchers 40%Predictive Machine Learning Practitioners 35%Applied Statisticians 25%
Causal Inference Researchers
Economists and epidemiologists who prioritize unbiased coefficient estimates over tight confidence intervals.
Predictive Machine Learning Practitioners
Data scientists focused entirely on out-of-sample forecasting accuracy rather than causal explanation.
Applied Statisticians
Pragmatic analysts who seek a middle ground through dimensionality reduction and composite indices.

Perspectives this story doesn't cover

  • Software developers building automated feature-selection algorithms

Sources

Source coverage

6 outlets

3 viewpoints surfaced

Causal Inference Researchers 40%Predictive Machine Learning Practitioners 35%Applied Statisticians 25%
  1. [1]Wikipedia

    Omitted-variable bias

    Read on Wikipedia →
  2. [2]Wikipedia

    Variance inflation factor

    Read on Wikipedia →
  3. [3]Wikipedia

    Gauss–Markov theorem

    Read on Wikipedia →
  4. [4]Penn State Eberly College of ScienceApplied Statisticians

    10.4 - Multicollinearity

    Read on Penn State Eberly College of Science →
  5. [5]Scribbr

    Omitted Variable Bias | Definition, Examples & Causes

    Read on Scribbr →
  6. [6]Factlen Editorial TeamCausal Inference Researchers

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team →

Comments

Stay informed

Every angle. Every day.

Get Data & Analysis stories with full source coverage and perspective breakdowns, free every day.