P-Value vs. Confidence Interval: Why the Range of Plausible Effect Sizes is More Informative Than a Significance Threshold
While a p-value only indicates if an effect exists mathematically, a confidence interval reveals the actual magnitude and precision of that effect. This shift in statistical reporting helps researchers and readers distinguish between a technically significant finding and a practically useful one.
By Paige Carter
- Statistical Reformers
- Advocates pushing to abandon binary significance testing entirely in favor of estimation.
- Clinical Practitioners
- Medical professionals who prioritize the physical magnitude of a treatment's effect over its mathematical probability.
- Traditional Methodologists
- Statisticians who defend the p-value as a valid metric when interpreted correctly as continuous evidence.
Perspectives this story doesn't cover
- Journal Editors
- Peer Reviewers
At a glance
- A p-value only measures the probability of data under a null hypothesis, not the size or importance of an effect.
- Massive sample sizes can generate statistically significant p-values for microscopic, practically useless results.
- Confidence intervals provide a point estimate and a margin of error, showing the actual range of plausible outcomes.
- Major statistical bodies and medical journals are actively pushing researchers to report effect sizes rather than binary significance thresholds.
A medical trial might report a statistically significant drop in blood pressure, but the 95% confidence interval reveals the actual magnitude: a reduction of anywhere from 0.1 to 0.3 millimeters of mercury. That physical basis—the tangible measurement of the effect rather than a binary mathematical flag—is why statistical bodies are pushing to replace the traditional p-value. Researchers have historically relied on a strict threshold to determine if an experiment worked, but modern guidelines demand that scientists show the full range of plausible outcomes.[9]
The core problem with the p-value is that it measures the wrong thing for practical decision-making. It calculates the probability of seeing the observed data if the null hypothesis—the assumption that there is zero effect—were perfectly true. It does not measure the size, importance, or clinical relevance of the finding. As the American Statistical Association explicitly warned in a landmark 2016 statement, "Scientific conclusions and business or policy decisions should not be based only on whether a p-value passes a specific threshold."[1][4]
This reliance on a strict cutoff, typically a p-value of less than 0.05, creates a distorted incentive structure. A study with a massive sample size of 50,000 patients can detect a microscopic, clinically useless difference and still achieve a highly significant p-value of 0.001. Conversely, a smaller study might uncover a massive, life-saving effect but fail to cross the 0.05 threshold due to limited data, causing a genuinely useful intervention to be discarded as a failure.[5][7]
Confidence intervals solve this by shifting the focus from hypothesis testing to estimation. A 95% confidence interval provides a point estimate—the most likely effect size—surrounded by a margin of error. If a new weight-loss drug shows an average loss of 10 pounds, the confidence interval might reveal that the true effect in the broader population likely falls between 8 and 12 pounds. This gives doctors and patients a concrete range to evaluate whether the treatment is worth the cost and side effects.[3][6]
The distinction between statistical significance and clinical importance is starkest in medical research. The journal Anesthesia & Analgesia highlighted this in 2018, noting that a statistically significant result simply means the effect is unlikely to be exactly zero. However, if a pain medication reduces a patient's discomfort by 0.2 points on a 10-point scale, the fact that the result is statistically significant does not make the drug clinically useful to the patient experiencing the pain.[4]
The push for estimation over testing is not new. The British Medical Journal began advocating for confidence intervals over p-values as early as 1986, arguing that the medical literature was plagued by a binary approach to data. Despite this decades-old guidance, the entrenchment of the 0.05 threshold in statistical software and peer-review standards kept the p-value dominant across the sciences.[6][8]
The breaking point arrived when multiple disciplines recognized a widespread replication crisis. In 2019, researchers writing in Biology Letters declared that "the reign of the p-value is over," asking what alternative analyses could fill the power vacuum. The consensus across biology, medicine, and psychology converged on reporting effect sizes alongside confidence intervals, forcing authors to defend the magnitude of their findings rather than hiding behind a single probability metric.[8]
The breaking point arrived when multiple disciplines recognized a widespread replication crisis.
To clarify the path forward, the American Statistical Association convened a President's Task Force in 2021 to issue a definitive statement on statistical significance and replicability. The task force concluded that while p-values can still serve as a continuous measure of evidence against a null hypothesis, they should never be used to make dichotomous claims of "significant" or "not significant."[2]
The mechanics of a confidence interval also provide a built-in warning system for uncertainty. A wide interval—such as a projected revenue increase of anywhere from $10,000 to $900,000—immediately signals to a business leader that the data is too noisy to support a firm decision. A p-value of 0.04 on that same dataset would simply flash a green light, masking the extreme variance underlying the average.[7][9]
Furthermore, confidence intervals make it immediately obvious when a study's results are inconclusive. If a 95% confidence interval for a difference in heart attack rates spans from a 5% reduction to a 2% increase, the interval crosses zero. This visually and mathematically demonstrates that the data cannot rule out the possibility that the treatment actually causes harm, providing a more intuitive grasp of risk than a p-value of 0.15.[5]
Transitioning away from p-values requires a fundamental change in how students are taught and how software defaults are configured. Most introductory statistics courses still spend weeks on null-hypothesis significance testing before briefly introducing effect sizes at the end of the semester. Reversing this order—teaching estimation first—is now a primary goal of statistical educators.[1][9]
For readers evaluating scientific claims, the actionable takeaway is to ignore the word "significant" and look for the physical units. Whether evaluating a new battery technology's cycle life, a dietary supplement's impact on cholesterol, or a marketing campaign's conversion rate, demanding the confidence interval ensures that the focus remains on the actual, real-world magnitude of the result.[3]
The transition is already visible in top-tier medical journals. Editors now routinely reject manuscripts that report p-values without accompanying confidence intervals and effect sizes. This editorial enforcement is proving more effective than decades of academic debate, as researchers are forced to adapt their reporting to get published.[4][5]
Even in fields outside of medicine, such as economics and engineering, the shift is taking hold. When an engineer tests the tensile strength of a new alloy, knowing that it is "significantly stronger" than steel is useless for designing a bridge. The engineer needs the confidence interval of the yield strength to calculate safety margins.[9]
The goal of this statistical reform is not to banish the p-value entirely, but to strip it of its status as the sole arbiter of scientific truth. By pairing probability metrics with confidence intervals, researchers provide a complete picture: the likelihood that an effect exists, combined with a realistic estimate of how large that effect might be. The next verifiable checkpoint for this shift will be the complete removal of the term "statistically significant" from the author guidelines of major scientific publishers.[2]
Terms to know
- P-value
- The probability of obtaining the observed results if the assumption that there is no actual effect (the null hypothesis) is completely true.
- Confidence Interval (CI)
- A range of values that is likely to contain the true effect size, providing both the most likely estimate and the margin of error.
- Effect Size
- The actual magnitude or physical size of the difference or relationship measured in a study.
- Null Hypothesis
- The default assumption in an experiment that there is no relationship, no difference, or no effect.
- Statistical Significance
- A mathematical determination that the observed data is unlikely to have occurred by random chance, typically declared when a p-value falls below 0.05.
Questions readers ask
Why is a p-value of 0.05 the standard cutoff?
The 0.05 threshold is largely a historical accident, popularized by statistician Ronald Fisher in the 1920s as a convenient rule of thumb, not a hard mathematical law.
Can a result be statistically significant but practically useless?
Yes. In studies with very large sample sizes, even microscopic and clinically irrelevant differences can produce highly significant p-values.
What does it mean if a confidence interval crosses zero?
If a confidence interval for a difference includes zero, it means the data cannot rule out the possibility that there is no effect, or even a negative effect.
Sources
[1]The American StatisticianStatistical ReformersThe ASA's Statement on p-Values: Context, Process, and Purpose
Read on The American Statistician →
[2]The Annals of Applied StatisticsTraditional MethodologistsASA President's Task Force Statement on Statistical Significance and Replicability
Read on The Annals of Applied Statistics →
[3]Journal of Clinical HypertensionClinical PractitionersIn Praise of Confidence Intervals: Much More Informative Than P Values Alone
Read on Journal of Clinical Hypertension →
[4]Anesthesia & AnalgesiaClinical PractitionersStatistical Significance Versus Clinical Importance of Observed Effect Sizes: What Do P Values and Confidence Intervals Really Represent?
Read on Anesthesia & Analgesia →
[5]Korean Journal of AnesthesiologyAlternatives to P value: confidence interval and effect size
Read on Korean Journal of Anesthesiology →
[6]BMJClinical PractitionersConfidence intervals rather than P values: estimation rather than hypothesis testing.
Read on BMJ →
[7]What is...? seriesWhat are confidence intervals and p-values?
Read on What is...? series →
[8]Biology LettersStatistical ReformersThe reign of the p-value is over: what alternative analyses could we employ to fill the power vacuum?
Read on Biology Letters →
[9]Factlen Editorial TeamSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
Comments
More in Guides
See all →Tire Ratings
Treadwear, Traction, and Temperature Ratings: How the UTQG System Predicts Tire Longevity and Wet Grip
5 sources
Windows Security
How Windows Secretly Tags Every Downloaded File to Trigger Security Sandboxes
4 sources
Research Methods
How Randomization and Control Groups Isolate Causation from Correlation in Scientific Studies
6 sources
HVAC Upgrades
Smart vs. Programmable Thermostats: The True ROI and Energy Savings Comparison
3 sources
Every angle. Every day.
Get Guides stories with full source coverage and perspective breakdowns delivered to your inbox.




