Skip to main content
ExplainerStatistical MethodologyData Science· 7 min read· in Data & Analysis

Dichotomizing Continuous Predictors Discards 36 Percent of Information: Why Median Splits Distort Estimates

Converting continuous variables into binary categories destroys statistical power and actively manufactures false positive results. Despite decades of methodological warnings, the median split remains a pervasive analytical error across scientific disciplines.

By Nicolas Laurent

In short

  • Splitting a continuous variable at its median mathematically destroys 36 percent of the dataset's variance and statistical power.
  • Categorizing correlated predictors actively manufactures false-positive interactions, creating the illusion of scientific discovery where none exists.
  • Modern statistical software provides robust alternatives, such as restricted cubic splines, that model non-linear relationships without discarding data.

A clinical trial collects exact blood pressure readings for 10,000 patients to test a new cardiovascular intervention. Before running the analysis, the lead researcher draws a line down the middle, labeling half the patients "high" and half "low."

In a single keystroke, the study has just thrown away the statistical equivalent of 3,600 patients. This practice, known as dichotomization or a median split, remains one of the most common and damaging habits in modern data analysis.

By converting a continuous variable into a binary category, analysts intentionally erase the nuance between a patient just above the median and one at the extreme high end. The statistical penalty for this convenience is severe and mathematically unavoidable.

Decades of methodological research confirm that splitting a continuous predictor at its median discards approximately 36 percent of the information contained in that variable. The loss of variance fundamentally degrades the model's ability to detect true relationships.[1][9]

"Dichotomising a continuous variable is equivalent to discarding a third of the data," researchers noted in a landmark methodological review for The BMJ. The practice artificially compresses natural variation into arbitrary, homogeneous blocks.[1]

The Mathematics of Information Loss

To understand the 36 percent penalty, consider the mechanics of statistical variance. Variance is the mathematical engine that drives discovery; it is the signal that separates a true effect from random background noise in any observational dataset.

A median split mathematically destroys 36 percent of the Fisher information contained in a continuous variable.

Dichotomization deliberately crushes this variance, pulling every data point toward the artificial center of its new category. A 40-year-old and an 80-year-old are mathematically treated as identical if both happen to fall into the "older" demographic bucket.

When the variance is compressed, the standard error of the estimate inflates dramatically. To achieve the same statistical power as an analysis using the continuous variable, a study using a median split must increase its sample size by roughly 50 percent.[1]

For a clinical trial costing thousands of dollars per participant, this represents a massive, unforced financial and scientific error. The loss of power directly increases the risk of Type II errors, where true effects are missed entirely.[2]

Manufacturing False Interactions

The damage of dichotomization extends far beyond merely missing true effects through reduced power. Under specific conditions, categorizing continuous variables actively manufactures false positives, creating the illusion of scientific discovery where none exists.[5]

This phenomenon is particularly severe in multiple regression models where two or more continuous predictors are correlated. When one of these correlated predictors is dichotomized, the residual variance does not simply disappear from the mathematical environment.[9]

Instead, that leftover information bleeds into the other variables in the model. This leakage can create a spurious statistical interaction, leading researchers to falsely conclude that an effect only occurs under specific, categorized conditions.[5]

Methodologists in consumer psychology demonstrated this Type I error inflation in a 2015 ScienceDirect analysis. They showed that median splits routinely generate false-positive interactions, polluting the psychological literature with effects that cannot be replicated.[8]

Categorizing correlated predictors forces residual variance into other variables, manufacturing false-positive interactions.

The Columbia University statistical modeling group has noted that if a split is absolutely necessary for presentation purposes, dividing the data into three parts is significantly less destructive than a median split.[6]

By comparing the top third to the bottom third and discarding the middle, researchers can sometimes isolate extreme effects without manufacturing the same degree of false interactions. However, keeping the variable continuous remains the optimal choice.[6]

Clinical and Ecological Consequences

The consequences of these statistical distortions are highly visible across applied scientific disciplines. In neuroradiology, researchers analyzing the natural history of unruptured brain aneurysms found that categorizing continuous variables led to fundamentally flawed risk assessments.[3]

By grouping aneurysm sizes into arbitrary bins, the clinical models failed to capture the exponential increase in rupture risk as the physical size increased. The binary model simply viewed a wide range of dangerous aneurysms as identically "large."[3]

The American Journal of Neuroradiology explicitly warned that analysis by categorizing continuous variables is "inadvisable." They noted that a patient with a 9-millimeter aneurysm might be placed in the same risk category as one with a 5-millimeter aneurysm.[3]

Ecology and evolutionary biology suffer from similar methodological blind spots. A review in the Proceedings of the Royal Society B highlighted how categorizing continuous traits like body mass or temperature obscures the subtle gradients that drive natural selection.[4]

When ecologists force continuous biological variation into discrete bins, they risk misinterpreting the fundamental drivers of survival. A species' response to climate change is rarely a binary switch; it is a continuous curve that categorization completely fails to capture.[4]

The Illusion of Clinical Utility

The primary defense of dichotomization is clinical or practical utility. Physicians do not prescribe medication on a continuous curve; they make binary decisions to treat or not to treat based on established diagnostic thresholds.

Binary models fail to capture exponential risk increases, treating vastly different data points as identical.

Therefore, the argument goes, the statistical models should reflect these binary decision points from the beginning. This defense conflates the goal of statistical modeling with the entirely separate goal of clinical application and treatment delivery.[10]

The purpose of a statistical model is to accurately describe the relationship between variables in the natural world. Only after that relationship is accurately mapped using all available continuous data should clinical action thresholds be applied.[10]

Forcing the data into a binary format before the analysis even begins guarantees that the resulting model will be inaccurate. A continuous model can always generate binary predictions later, but a dichotomized model can never recover the lost information.

Furthermore, clinical thresholds are rarely static. The definition of high blood pressure or abnormal cholesterol has shifted repeatedly over the past three decades as new evidence emerges, rendering models built on historical cut-points instantly obsolete.

When researchers lock their analysis to a specific threshold, they tie their findings to the medical consensus of that exact year. A continuous model, by contrast, provides a timeless mathematical equation that future researchers can adapt to new guidelines.

Better Alternatives for Complex Data

Modern statistical software provides numerous alternatives to dichotomization that preserve the integrity of continuous data. When the relationship between a predictor and an outcome is strictly linear, the continuous variable should simply be left in its original state.[7]

In cases where the relationship is non-linear—such as a drug that is beneficial at moderate doses but toxic at high doses—researchers can utilize fractional polynomials. These techniques allow the model to fit a flexible curve to the data.[1]

Restricted cubic splines allow models to fit non-linear biological realities without discarding continuous data.

Another robust alternative is the use of restricted cubic splines. Splines work by dividing the continuous variable into segments at specific points called knots, and fitting a separate, smooth curve to each individual segment.

Crucially, these curves are forced to join smoothly at the knots, creating a continuous, nuanced map of the variable's effect across its entire range. They retain 100 percent of the collected information while capturing complex biological realities.

While these methods require slightly more statistical expertise to implement, they are now standard features in all major analytical software packages. The historical excuse that continuous non-linear modeling is too computationally difficult is no longer valid.

The Path Forward for Data Science

Eradicating the median split requires a cultural shift in how research is reviewed and published. Peer reviewers and journal editors must begin treating the unjustified dichotomization of continuous variables as a fatal methodological flaw.

Some medical and statistical journals have already taken this step. They explicitly warn authors that analyses relying on categorized continuous variables will be rejected unless a compelling, pre-specified biological justification is provided for the threshold.[3][5]

Educational curricula in the social and biological sciences must also adapt. Many introductory statistics courses still teach median splits as a valid way to simplify complex data for basic variance testing, inadvertently training researchers to destroy their own data.

The replication crisis in psychology and medicine has forced a reckoning with sloppy statistical habits. As disciplines tighten their standards for evidence, the tolerance for unforced errors like the median split is rapidly evaporating across the scientific community.

The replication crisis in psychology and medicine has forced a reckoning with sloppy statistical habits.

Ultimately, the integrity of observational research depends on respecting the data as it was collected. Nature rarely operates in binary switches; it exists in gradients, spectrums, and continuous curves that demand high-resolution analytical approaches.

Every time a continuous variable is split at the median, a third of the hard-won data is discarded, and the risk of publishing a false conclusion rises. Researchers can no longer afford to throw away 36 percent of their information just to make the math look simpler.

How we did this

Method
Synthesizing the mathematical penalty of dichotomization across distinct disciplines (medicine, ecology, consumer psychology) to quantify the universal statistical cost of median splits on statistical power and false positive rates.
What we found
The practice of median splitting not only universally degrades statistical power equivalent to discarding over a third of a dataset, but actively manufactures false positive interactions regardless of the scientific discipline, making it a net-negative analytical choice in all observational research.
What we worked from
Limits of this analysis
This analysis applies to continuous variables with linear or monotonic relationships; threshold effects where a true biological or physical step-change exists may legitimately require categorization.

Definitions

Dichotomization
The process of converting a continuous variable into two distinct categories, usually by splitting the data at a specific threshold.
Median Split
A specific form of dichotomization where data is divided exactly in half, creating a high and low group based on the median value.
Type I Error
A false positive result, where a statistical test incorrectly indicates a significant effect or interaction that does not actually exist.
Type II Error
A false negative result, where a statistical test fails to detect a true effect because the model lacks sufficient power.
Variance
A mathematical measurement of how far individual data points are spread out from their average value, essential for detecting statistical signals.
Restricted Cubic Spline
A mathematical technique that fits flexible, continuous curves to data segments, allowing for non-linear modeling without categorization.

Questions & answers

Does dichotomization ever make statistical sense?

Only when the underlying phenomenon is genuinely binary in nature, such as a biological threshold where an effect only triggers above a specific, absolute concentration.

Why is the information loss exactly 36 percent?

It is mathematically derived from the loss of Fisher information when a normally distributed continuous variable is reduced to two categories at its median.

Can researchers just increase sample size to compensate?

While increasing the sample size by roughly 50 percent restores statistical power, it does not fix the inflated risk of manufacturing false-positive interactions.

Are tertiles or quartiles better than median splits?

They are slightly less destructive to variance than a two-way split, but they still discard significant information compared to keeping the variable continuous.

Analysis by camp

Methodologists and Statisticians

Argue that dichotomization is mathematically indefensible and actively harms the scientific record.

Statistical methodologists view the median split as an unforced error that degrades the quality of observational research. By intentionally discarding 36 percent of a dataset's variance, researchers artificially inflate their standard errors and dramatically increase the risk of Type II errors. More concerning to methodologists is the inflation of Type I errors; when correlated continuous predictors are categorized, the residual variance bleeds into interaction terms, manufacturing false-positive findings that pollute the scientific literature and cannot be replicated by independent teams.

Clinical and Applied Researchers

Often defend categorization as a necessary bridge between abstract mathematics and real-world clinical decision-making.

Applied researchers, particularly in medicine and public health, frequently argue that statistical models must reflect the reality of treatment delivery. Physicians do not prescribe interventions on a continuous curve; they rely on binary diagnostic thresholds to make immediate decisions. From this perspective, categorizing a continuous biomarker into "high" and "normal" risk buckets makes the resulting statistical model immediately actionable for frontline healthcare workers, even if it sacrifices a degree of mathematical precision in the process.

Statistical Software Developers

Focus on building accessible tools to remove the computational friction that historically drove researchers toward simple median splits.

Software developers and data scientists recognize that the historical reliance on median splits was largely driven by computational limitations and the difficulty of modeling non-linear relationships in early statistical packages. Their solution has been to democratize advanced techniques like restricted cubic splines and generalized additive models (GAMs). By making these tools standard, accessible features in modern software, developers aim to eliminate the practical excuses for dichotomization, allowing researchers to fit flexible, continuous curves to complex biological realities without requiring advanced programming skills.

Methodologists and Statisticians 50%Clinical and Applied Researchers 30%Ecologists and Psychologists 20%
Methodologists and Statisticians
Argue that dichotomization is mathematically indefensible and actively harms the scientific record by inflating false positives.
Clinical and Applied Researchers
Often defend categorization as a necessary bridge between abstract mathematics and real-world clinical decision-making.
Ecologists and Psychologists
Highlight how arbitrary data splits obscure the subtle gradients that drive natural selection and human behavior.

Perspectives this story doesn't cover

  • Journal peer reviewers
  • Undergraduate statistics educators

Sources

Source coverage

10 outlets

3 viewpoints surfaced

Methodologists and Statisticians 50%Clinical and Applied Researchers 30%Ecologists and Psychologists 20%
  1. [1]The BMJMethodologists and Statisticians

    The cost of dichotomising continuous variables

    Read on The BMJ →
  2. [2]PubMedClinical and Applied Researchers

    Consequences of dichotomization

    Read on PubMed →
  3. [3]American Journal of NeuroradiologyClinical and Applied Researchers

    Analysis by Categorizing or Dichotomizing Continuous Variables Is Inadvisable: An Example from the Natural History of Unruptured Aneurysms

    Read on American Journal of Neuroradiology →
  4. [4]Proceedings of the Royal Society B: Biological SciencesEcologists and Psychologists

    Overcoming the pitfalls of categorizing continuous variables in ecology, evolution and behaviour

    Read on Proceedings of the Royal Society B: Biological Sciences →
  5. [5]BMC Medical Research MethodologyMethodologists and Statisticians

    Spurious interaction as a result of categorization

    Read on BMC Medical Research Methodology →
  6. [6]Statistical Modeling, Causal Inference, and Social ScienceMethodologists and Statisticians

    Beyond the median split: Splitting a predictor into 3 parts

    Read on Statistical Modeling, Causal Inference, and Social Science →
  7. [7]The Analysis FactorMethodologists and Statisticians

    Continuous and Categorical Variables: The Trouble with Median Splits

    Read on The Analysis Factor →
  8. [8]ScienceDirectEcologists and Psychologists

    Median splits, Type II errors, and false–positive consumer psychology: Don't fight the power

    Read on ScienceDirect →
  9. [9]PubMedClinical and Applied Researchers

    Dichotomizing continuous predictors in multiple regression: a bad idea

    Read on PubMed →
  10. [10]Factlen Editorial Team

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team →

Comments

Stay informed

Every angle. Every day.

Get Data & Analysis stories with full source coverage and perspective breakdowns, free every day.