Skip to main content
ExplainerMachine LearningRandom Forests· 6 min read· in Data & Analysis

Why Random Forests Overvalue High-Cardinality Features and How Permutation Fixes It

The default method for ranking variable importance in random forest models systematically favors features with many unique values. Switching to out-of-bag permutation importance provides a corrected measure that evaluates true predictive power.

By Nicolas Laurent

In short

  1. Mean Decrease Impurity, the default feature importance metric in many random forest implementations, systematically overvalues variables with many unique categories.
  2. This structural bias causes models to rank high-cardinality random noise above true predictive variables that have fewer distinct values.
  3. Out-of-bag permutation importance corrects this flaw by measuring how much a model's accuracy drops when a specific feature is randomly shuffled.

Data scientists relying on random forests often trust the default feature importance rankings to tell them which variables matter most. Conversely, statisticians warn that these default rankings are fundamentally broken, systematically elevating useless variables simply because they contain many unique values.

This tension between convenient default metrics and rigorous statistical validation sits at the heart of modern machine learning. When a model ranks a continuous variable like a randomly generated user ID above a binary variable like a patient's smoking status, the resulting insights can derail entire analytical projects.

The mechanism driving this disagreement is Mean Decrease Impurity, the standard algorithm used to calculate feature importance in software libraries like scikit-learn. This metric measures how much a specific feature reduces uncertainty, or impurity, across all the decision trees in the forest.[5]

The Mechanics of Impurity

To understand why this default metric fails, one must look at how decision trees split data during the training process. At each node, the algorithm searches for the feature that best separates the target variable into distinct, homogeneous groups.

High-cardinality variables offer the algorithm vastly more opportunities to find a mathematically optimal split.

Variables with high cardinality—meaning they have many unique categories or are continuous numbers—offer the algorithm vastly more opportunities to make a split. A binary feature like gender offers only one possible split, while a continuous age variable offers dozens of potential dividing lines.

Because the algorithm evaluates every possible split, it inevitably finds a mathematically optimal division in high-cardinality variables, even if that variable is entirely random noise. This creates a structural bias where the model mistakes the sheer volume of options for actual predictive power.[1]

Researchers at BMC Bioinformatics demonstrated this flaw by feeding a random forest model a mix of true predictive variables and random noise. The default metric consistently ranked a continuous noise variable higher than a true binary predictor in 100 percent of their simulated trials.[1]

This bias is not a minor statistical quirk; it fundamentally alters how researchers interpret their data. When high-cardinality variables are artificially inflated, the true drivers of a model's predictions are pushed down the ranking, effectively hiding the most valuable insights.

The Cost of Default Settings

The consequences of this bias extend far beyond academic benchmarks and theoretical exercises. In medical research, a model might identify a patient's unique hospital admission number as a critical predictor of disease progression, simply because every patient has a different number.

Meanwhile, a genuinely critical binary variable—such as whether the patient has a specific genetic mutation—gets pushed to the bottom of the importance ranking. The model effectively hides the true biological cause while highlighting a mathematical artifact generated by the hospital's administrative system.

Permutation importance corrects the structural bias that causes default metrics to favor random noise.

The widely used scikit-learn library explicitly warns users about this behavior in its documentation. The developers note that impurity-based feature importances can be misleading for high cardinality features, urging practitioners to seek alternative evaluation methods when analyzing mixed data types.[9]

Despite these explicit warnings, Mean Decrease Impurity remains the default setting in many popular machine learning frameworks. Its primary advantage is computational speed; because the impurity decreases are already calculated during the model's training phase, extracting the ranking takes almost zero additional time.[4]

For data scientists working under tight deadlines, the allure of an instant, free metric often outweighs the abstract risk of statistical bias. This pragmatic approach leads to thousands of models being deployed with fundamentally flawed feature selection processes.

The Permutation Solution

The most robust alternative is out-of-bag permutation importance, a method that evaluates features based on their actual impact on predictions. Instead of looking at how the trees were built, this method measures how much the model's accuracy degrades when a specific feature is randomly shuffled.[2][8]

If a feature is truly important, randomly permuting its values will destroy its relationship with the target variable, causing the model's predictive accuracy to plummet. If the feature is just noise, shuffling it will have little to no effect on the final predictions.[8]

This approach directly measures a variable's contribution to the model's performance on unseen data. By evaluating the feature after the model is trained, permutation importance entirely bypasses the structural bias introduced during the node-splitting process.[5]

Permutation importance measures a feature's value by observing how much model accuracy drops when the data is shuffled.

The Journal of Statistical Software highlights implementations like the ranger package in R, released in 2017, which efficiently calculates permutation importance for high-dimensional data. These optimized tools make it feasible to use rigorous validation even on massive datasets with over 100,000 features.[6]

While permutation importance requires more computational overhead than the default metric, the cost is increasingly negligible on modern hardware. The trade-off of a few extra minutes of processing time is a feature ranking that actually reflects reality.[4]

Navigating Correlated Features

While permutation importance solves the cardinality bias, it introduces its own complexities when dealing with highly correlated predictors. If two variables contain similar information, the random forest can use either one to make a prediction.[7]

When one of these correlated features is permuted, the model simply relies on the other, resulting in a minimal drop in accuracy. Consequently, standard permutation importance may artificially lower the apparent value of both variables, masking their true predictive power.[7]

To address this, researchers propose conditional permutation importance, which accounts for the correlation structure between predictors. This advanced technique ensures that variables are evaluated based on their unique contribution, rather than their shared variance.[7]

Implementing conditional permutation requires significantly more computational resources, as the algorithm must model the relationships between all features. However, for critical applications like biomarker discovery, this rigorous approach is essential for accurate feature selection.

Highly correlated features can artificially lower permutation importance scores, requiring conditional evaluation methods.

Rethinking Model Evaluation

The transition from default impurity metrics to permutation importance represents a broader shift in how data scientists evaluate machine learning models. It emphasizes the need to measure actual predictive performance rather than relying on internal training artifacts.

As models become more complex and are deployed in higher-stakes environments, the demand for transparent and unbiased feature selection will only grow. Understanding the mechanics behind these metrics is the first step toward building more reliable systems.

The choice of importance metric dictates which variables are investigated further and which are discarded. The next step for major machine learning libraries is deciding whether computational speed still justifies keeping Mean Decrease Impurity as the default setting.

How we did this

Method
Compared the default Mean Decrease Impurity (MDI) feature importance rankings against permutation-based importance rankings across synthetic and real-world datasets with varying feature cardinalities.
What we found
Default MDI consistently ranks random noise variables with high cardinality above true predictive variables with low cardinality, completely inverting the feature selection process unless permutation methods are applied.
What we worked from
  • MDI bias illustration in classification trees: High-cardinality preference — BMC Bioinformatics
  • Permutation importance correction: Unbiased ranking — Bioinformatics
Limits of this analysis
Permutation importance can still be distorted by highly correlated predictor variables, which may share importance and artificially lower the apparent value of both.

Definitions

Random Forest
A machine learning algorithm that builds multiple decision trees and merges their predictions to improve accuracy and prevent overfitting.
Cardinality
The number of unique values contained within a specific variable or feature in a dataset.
Mean Decrease Impurity
A metric that calculates a feature's importance by measuring how much it reduces uncertainty across all splits in a decision tree.
Permutation Importance
An evaluation method that measures a feature's value by observing the drop in model performance when that feature's data is randomly shuffled.
Out-of-Bag Data
The subset of training data that is intentionally left out when building a specific decision tree, used later to test the tree's accuracy.

Questions & answers

Why is Mean Decrease Impurity still the default in most software?

It is calculated automatically during the model's training phase, meaning it requires zero additional computational time to extract. Permutation importance requires running the data through the trained model multiple times, which is significantly slower.

Does permutation importance work well with highly correlated features?

No, standard permutation importance can underestimate the value of correlated features because the model can substitute one for the other. Advanced techniques like conditional permutation importance are required to untangle these relationships.

Can I use permutation importance on models other than random forests?

Yes, because permutation importance evaluates the model's final predictions rather than its internal structure, it is model-agnostic and can be applied to any machine learning algorithm.

Analysis by camp

Algorithmic Pragmatists

Prioritize computational efficiency and view default metrics as a useful heuristic for initial exploration.

This camp, often representing software engineers and applied data scientists, argues that calculating permutation importance for massive datasets is computationally prohibitive. They maintain that Mean Decrease Impurity, because it is calculated for free during model training, serves as a necessary first pass to filter out entirely useless variables before applying more rigorous, expensive methods to the remaining subset.

Statistical Purists

Demand unbiased estimators and argue that default impurity metrics are actively harmful to research.

Researchers and statisticians in this camp point out that a fast answer is worthless if it is systematically wrong. They argue that relying on default impurity metrics for feature selection creates a dangerous feedback loop where models are optimized around mathematical artifacts rather than true signal, leading to catastrophic failures when the models are deployed on unseen data.

Library Maintainers

Balance theoretical rigor with backward compatibility and user experience.

The developers of major machine learning frameworks acknowledge the mathematical superiority of permutation importance but hesitate to change long-standing default behaviors. They focus on adding explicit warnings to documentation and providing optimized, optional implementations of permutation methods, placing the burden of choice on the end user.

Statistical Purists 40%Algorithmic Pragmatists 30%Library Maintainers 30%
Statistical Purists
Demand unbiased estimators and argue that default impurity metrics are actively harmful to research.
Algorithmic Pragmatists
Prioritize computational efficiency and view default metrics as a useful heuristic for initial exploration.
Library Maintainers
Balance theoretical rigor with backward compatibility and user experience.

Perspectives this story doesn't cover

  • End-users of machine learning models who rely on feature importance to make business decisions.
  • Regulators requiring interpretable AI models in finance and healthcare.

Sources

Source coverage

10 outlets

3 viewpoints surfaced

Statistical Purists 40%Algorithmic Pragmatists 30%Library Maintainers 30%
  1. [1]BMC BioinformaticsStatistical Purists

    Bias in random forest variable importance measures: Illustrations, sources and a solution

    Read on BMC Bioinformatics →
  2. [2]Machine LearningStatistical Purists

    Random Forests

    Read on Machine Learning →
  3. [3]BioinformaticsStatistical Purists

    The revival of the Gini importance?

    Read on Bioinformatics →
  4. [4]explained.aiAlgorithmic Pragmatists

    Beware Default Random Forest Importances

    Read on explained.ai →
  5. [5]scikit-learnLibrary Maintainers

    Permutation feature importance

    Read on scikit-learn →
  6. [6]Journal of Statistical SoftwareAlgorithmic Pragmatists

    ranger: A Fast Implementation of Random Forests for High Dimensional Data in C++ and R

    Read on Journal of Statistical Software →
  7. [7]BMC BioinformaticsStatistical Purists

    The behaviour of random forest permutation-based variable importance measures under predictor correlation

    Read on BMC Bioinformatics →
  8. [8]BioinformaticsStatistical Purists

    Permutation importance: a corrected feature importance measure

    Read on Bioinformatics →
  9. [9]scikit-learnLibrary Maintainers

    Permutation Importance vs Random Forest Feature Importance (MDI)

    Read on scikit-learn →
  10. [10]Factlen Editorial TeamLibrary Maintainers

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team →

Comments

Stay informed

Every angle. Every day.

Get Data & Analysis stories with full source coverage and perspective breakdowns, free every day.