Skip to main content
ExplainerStatistical MethodsExplainer· 6 min read· in Guides

P-Value vs. Confidence Interval: Why the Range of Plausible Effect Sizes is More Informative Than a Significance Threshold

While a p-value only indicates if an effect exists mathematically, a confidence interval reveals the actual magnitude and precision of that effect. This shift in statistical reporting helps researchers and readers distinguish between a technically significant finding and a practically useful one.

By Paige Carter

Statistical Reformers 45%Clinical Practitioners 35%Traditional Methodologists 20%
Statistical Reformers
Advocates pushing to abandon binary significance testing entirely in favor of estimation.
Clinical Practitioners
Medical professionals who prioritize the physical magnitude of a treatment's effect over its mathematical probability.
Traditional Methodologists
Statisticians who defend the p-value as a valid metric when interpreted correctly as continuous evidence.

Perspectives this story doesn't cover

  • Journal Editors
  • Peer Reviewers

At a glance

  1. A p-value only measures the probability of data under a null hypothesis, not the size or importance of an effect.
  2. Massive sample sizes can generate statistically significant p-values for microscopic, practically useless results.
  3. Confidence intervals provide a point estimate and a margin of error, showing the actual range of plausible outcomes.
  4. Major statistical bodies and medical journals are actively pushing researchers to report effect sizes rather than binary significance thresholds.

A medical trial might report a statistically significant drop in blood pressure, but the 95% confidence interval reveals the actual magnitude: a reduction of anywhere from 0.1 to 0.3 millimeters of mercury. That physical basis—the tangible measurement of the effect rather than a binary mathematical flag—is why statistical bodies are pushing to replace the traditional p-value. Researchers have historically relied on a strict threshold to determine if an experiment worked, but modern guidelines demand that scientists show the full range of plausible outcomes.[9]

The core problem with the p-value is that it measures the wrong thing for practical decision-making. It calculates the probability of seeing the observed data if the null hypothesis—the assumption that there is zero effect—were perfectly true. It does not measure the size, importance, or clinical relevance of the finding. As the American Statistical Association explicitly warned in a landmark 2016 statement, "Scientific conclusions and business or policy decisions should not be based only on whether a p-value passes a specific threshold."[1][4]

This reliance on a strict cutoff, typically a p-value of less than 0.05, creates a distorted incentive structure. A study with a massive sample size of 50,000 patients can detect a microscopic, clinically useless difference and still achieve a highly significant p-value of 0.001. Conversely, a smaller study might uncover a massive, life-saving effect but fail to cross the 0.05 threshold due to limited data, causing a genuinely useful intervention to be discarded as a failure.[5][7]

A p-value only provides a binary pass/fail, while a confidence interval reveals the actual range of plausible outcomes.

Confidence intervals solve this by shifting the focus from hypothesis testing to estimation. A 95% confidence interval provides a point estimate—the most likely effect size—surrounded by a margin of error. If a new weight-loss drug shows an average loss of 10 pounds, the confidence interval might reveal that the true effect in the broader population likely falls between 8 and 12 pounds. This gives doctors and patients a concrete range to evaluate whether the treatment is worth the cost and side effects.[3][6]

The distinction between statistical significance and clinical importance is starkest in medical research. The journal Anesthesia & Analgesia highlighted this in 2018, noting that a statistically significant result simply means the effect is unlikely to be exactly zero. However, if a pain medication reduces a patient's discomfort by 0.2 points on a 10-point scale, the fact that the result is statistically significant does not make the drug clinically useful to the patient experiencing the pain.[4]

The push for estimation over testing is not new. The British Medical Journal began advocating for confidence intervals over p-values as early as 1986, arguing that the medical literature was plagued by a binary approach to data. Despite this decades-old guidance, the entrenchment of the 0.05 threshold in statistical software and peer-review standards kept the p-value dominant across the sciences.[6][8]

Confidence intervals make it immediately obvious when a study's results are too noisy to be useful, even if the average effect appears large.

The breaking point arrived when multiple disciplines recognized a widespread replication crisis. In 2019, researchers writing in Biology Letters declared that "the reign of the p-value is over," asking what alternative analyses could fill the power vacuum. The consensus across biology, medicine, and psychology converged on reporting effect sizes alongside confidence intervals, forcing authors to defend the magnitude of their findings rather than hiding behind a single probability metric.[8]

The breaking point arrived when multiple disciplines recognized a widespread replication crisis.

To clarify the path forward, the American Statistical Association convened a President's Task Force in 2021 to issue a definitive statement on statistical significance and replicability. The task force concluded that while p-values can still serve as a continuous measure of evidence against a null hypothesis, they should never be used to make dichotomous claims of "significant" or "not significant."[2]

The mechanics of a confidence interval also provide a built-in warning system for uncertainty. A wide interval—such as a projected revenue increase of anywhere from $10,000 to $900,000—immediately signals to a business leader that the data is too noisy to support a firm decision. A p-value of 0.04 on that same dataset would simply flash a green light, masking the extreme variance underlying the average.[7][9]

Furthermore, confidence intervals make it immediately obvious when a study's results are inconclusive. If a 95% confidence interval for a difference in heart attack rates spans from a 5% reduction to a 2% increase, the interval crosses zero. This visually and mathematically demonstrates that the data cannot rule out the possibility that the treatment actually causes harm, providing a more intuitive grasp of risk than a p-value of 0.15.[5]

A 95% confidence interval provides the most likely estimate surrounded by a calculated margin of error.

Transitioning away from p-values requires a fundamental change in how students are taught and how software defaults are configured. Most introductory statistics courses still spend weeks on null-hypothesis significance testing before briefly introducing effect sizes at the end of the semester. Reversing this order—teaching estimation first—is now a primary goal of statistical educators.[1][9]

For readers evaluating scientific claims, the actionable takeaway is to ignore the word "significant" and look for the physical units. Whether evaluating a new battery technology's cycle life, a dietary supplement's impact on cholesterol, or a marketing campaign's conversion rate, demanding the confidence interval ensures that the focus remains on the actual, real-world magnitude of the result.[3]

The transition is already visible in top-tier medical journals. Editors now routinely reject manuscripts that report p-values without accompanying confidence intervals and effect sizes. This editorial enforcement is proving more effective than decades of academic debate, as researchers are forced to adapt their reporting to get published.[4][5]

Even in fields outside of medicine, such as economics and engineering, the shift is taking hold. When an engineer tests the tensile strength of a new alloy, knowing that it is "significantly stronger" than steel is useless for designing a bridge. The engineer needs the confidence interval of the yield strength to calculate safety margins.[9]

The goal of this statistical reform is not to banish the p-value entirely, but to strip it of its status as the sole arbiter of scientific truth. By pairing probability metrics with confidence intervals, researchers provide a complete picture: the likelihood that an effect exists, combined with a realistic estimate of how large that effect might be. The next verifiable checkpoint for this shift will be the complete removal of the term "statistically significant" from the author guidelines of major scientific publishers.[2]

Terms to know

P-value
The probability of obtaining the observed results if the assumption that there is no actual effect (the null hypothesis) is completely true.
Confidence Interval (CI)
A range of values that is likely to contain the true effect size, providing both the most likely estimate and the margin of error.
Effect Size
The actual magnitude or physical size of the difference or relationship measured in a study.
Null Hypothesis
The default assumption in an experiment that there is no relationship, no difference, or no effect.
Statistical Significance
A mathematical determination that the observed data is unlikely to have occurred by random chance, typically declared when a p-value falls below 0.05.

Questions readers ask

Why is a p-value of 0.05 the standard cutoff?

The 0.05 threshold is largely a historical accident, popularized by statistician Ronald Fisher in the 1920s as a convenient rule of thumb, not a hard mathematical law.

Can a result be statistically significant but practically useless?

Yes. In studies with very large sample sizes, even microscopic and clinically irrelevant differences can produce highly significant p-values.

What does it mean if a confidence interval crosses zero?

If a confidence interval for a difference includes zero, it means the data cannot rule out the possibility that there is no effect, or even a negative effect.

Sources

Source coverage

9 outlets

3 viewpoints surfaced

Statistical Reformers 45%Clinical Practitioners 35%Traditional Methodologists 20%
  1. [1]The American StatisticianStatistical Reformers

    The ASA's Statement on p-Values: Context, Process, and Purpose

    Read on The American Statistician
  2. [2]The Annals of Applied StatisticsTraditional Methodologists

    ASA President's Task Force Statement on Statistical Significance and Replicability

    Read on The Annals of Applied Statistics
  3. [3]Journal of Clinical HypertensionClinical Practitioners

    In Praise of Confidence Intervals: Much More Informative Than P Values Alone

    Read on Journal of Clinical Hypertension
  4. [4]Anesthesia & AnalgesiaClinical Practitioners

    Statistical Significance Versus Clinical Importance of Observed Effect Sizes: What Do P Values and Confidence Intervals Really Represent?

    Read on Anesthesia & Analgesia
  5. [5]Korean Journal of Anesthesiology

    Alternatives to P value: confidence interval and effect size

    Read on Korean Journal of Anesthesiology
  6. [6]BMJClinical Practitioners

    Confidence intervals rather than P values: estimation rather than hypothesis testing.

    Read on BMJ
  7. [7]What is...? series

    What are confidence intervals and p-values?

    Read on What is...? series
  8. [8]Biology LettersStatistical Reformers

    The reign of the p-value is over: what alternative analyses could we employ to fill the power vacuum?

    Read on Biology Letters
  9. [9]Factlen Editorial Team

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team

Comments

Stay informed

Every angle. Every day.

Get Guides stories with full source coverage and perspective breakdowns delivered to your inbox.