Skip to main content
ExplainerStatistical MethodsLeague Tables· 8 min read· in Data & Analysis

Sampling Variance Scales Inversely With Size: Why Empirical Bayes Shrinkage Keeps Small Institutions From Dominating League Tables

When ranking institutions by performance, small sample sizes create statistical noise that artificially pushes small entities to the extreme top and bottom of league tables. Empirical Bayes shrinkage corrects this by pulling unreliable estimates toward the global average, ensuring rankings reflect true quality rather than random chance.

By Sofia Matos

In short

  1. Small sample sizes create high sampling variance, causing small institutions to artificially dominate the top and bottom of unadjusted rankings.
  2. Empirical Bayes shrinkage corrects this by pulling unreliable, low-sample estimates toward the global average.
  3. Major regulators now use these hierarchical models to ensure financial penalties target true systemic failures rather than statistical noise.

When a rookie baseball player steps up to the plate twice and hits the ball both times, their raw batting average is a perfect 1.000. If a sports analyst ranked all players by this unadjusted metric, that rookie would instantly displace every Hall of Fame veteran to become the greatest hitter in the history of the sport.[1]

The same mathematical illusion plagues healthcare, education, and public policy. A rural hospital that performs two specialized surgeries in a year, with one patient dying, records a 50 percent mortality rate. In a national league table, that small clinic immediately falls to the very bottom, appearing exceptionally dangerous.[2]

These extremes are not driven by exceptional skill or catastrophic failure, but by the mathematics of small sample sizes. When denominators are tiny, a single random event triggers a massive swing in the resulting percentage. To correct this distortion, statisticians rely on a technique called Empirical Bayes shrinkage.

The illusion of small samples

The core issue is sampling variance, which scales inversely with the size of the institution. In any dataset measuring performance rates, the entities with the fewest observations will inherently display the widest spread of outcomes. This volatility guarantees that small institutions will dominate both the top and bottom of any unadjusted ranking.

Statistician David Robinson famously demonstrated this using the Lahman Baseball Database. When looking at all players, those with the highest and lowest career batting averages are universally players with only a handful of at-bats. The true talent of a player with 10,000 at-bats is known precisely, but a player with five at-bats is mostly a product of statistical noise.[1]

The funnel of variance shows how small sample sizes produce extreme high and low rates.

This phenomenon appears constantly in public health data. In a widely cited analysis of kidney cancer incidence, the highest rate in the country was found in Cass County, with 41.1 cases per 100,000 people. However, that alarming rate was based on just seven cases in a population of 17,000.[1]

If Cass County had recorded just one fewer diagnosis that year, its incidence rate would have plummeted to 35.2 per 100,000, dropping it out of the top tier entirely. The extreme ranking was a mirage created by a small denominator. Large counties, bound by the law of large numbers, rarely reach such extremes.[1]

Tossing many different coins with binomial sampling will result in a greater variance than tossing a single coin, Robinson notes in his work on the subject. When we rank institutions without accounting for this variance, we are often just sorting them by their sample sizes rather than their actual performance.[1]

The league table problem

When governments and regulators publish league tables, this statistical artifact has severe real-world consequences. A 2009 study by the National Institutes of Health examined hospital rankings for pancreatic resections. They found that a hospital performing 10 procedures with two deaths was routinely flagged as a severe outlier due to its 20 percent mortality rate.[2]

Because of the small caseload, there is a considerable likelihood that the 20 percent estimate is the result of random chance rather than systemic medical errors. Meanwhile, a massive urban hospital with 500 cases and a true mortality rate of 8 percent might escape scrutiny, safely hidden in the middle of the pack.[2]

This creates a structural penalty for small institutions. In pay-for-performance programs, where funding is tied to ranking positions, small hospitals face a dramatic inverse relationship between their size and the stability of their ranks. A few unlucky cases can trigger devastating financial penalties.[2]

Small institutions face massive ranking instability when evaluated purely on raw averages.

To solve this, analysts must find a way to separate the true institutional signal from the noise of sampling variance. One approach is to simply exclude small institutions from the rankings entirely, but this discards valuable data and leaves rural populations without any quality metrics.

Instead, modern statistical frameworks use shrinkage estimators. By acknowledging that a 20 percent mortality rate on 10 cases is highly uncertain, the math can adjust the estimate to reflect what we actually know. This is where the Empirical Bayes method transforms the analysis.[2]

Enter Empirical Bayes

Traditional Bayesian statistics requires the analyst to specify a prior distribution before looking at the data. This subjective choice can be controversial in public policy, where stakeholders demand objective, data-driven metrics. Empirical Bayes bypasses this controversy by learning the prior directly from the dataset itself.

Rather than guessing the distribution of hospital mortality rates, the Empirical Bayes method calculates the grand mean and the true variance across the entire ensemble of hospitals. It uses the collective experience of all institutions to establish a baseline expectation for any single facility.

Empirical Bayes refers to a hacker approach to a specific kind of Bayesian Analysis, Robinson explains. Because the prior is estimated from the data, it provides a principled, computationally efficient way to regularize group means without relying on arbitrary assumptions.[1]

Once the global distribution is established, the algorithm evaluates each individual institution. It asks a simple question: how much evidence does this specific hospital or school provide to convince us that it is genuinely different from the global average?

The shrinkage factor pulls unreliable raw estimates back toward the global baseline.

For a massive institution with thousands of records, the evidence is overwhelming, and its raw average is trusted. For a tiny institution with only a dozen records, the evidence is weak. The algorithm responds by pulling the weak estimate back toward the global baseline.

The mechanics of shrinkage

The mathematical engine driving this correction is the shrinkage factor. The Empirical Bayes estimate is essentially a weighted average of the institution's raw score and the global mean. The weight assigned to the raw score depends entirely on its reliability.

The formula calculates the shrinkage factor by dividing the sampling variance of the specific institution by the total variance in the dataset. If the sampling variance is high, because the denominator is small, the shrinkage factor approaches 1.0, meaning the estimate is heavily shrunk toward the grand mean.

Conversely, if the institution has a massive sample size, its sampling variance approaches zero. The shrinkage factor drops, and the institution's adjusted score remains almost identical to its raw score. The math automatically scales the correction to the exact level of uncertainty.

Returning to the NIH hospital study, the global average mortality rate for the procedure was 5 percent. For the small hospital with a 20 percent raw rate on 10 cases, the Empirical Bayes formula recognized the extreme sampling variance and shrunk the estimate aggressively.[2]

The adjusted mortality rate for that small hospital was pulled all the way down to roughly 6 percent. The math determined that the two deaths were likely a statistical blip, and that the hospital's true underlying quality was much closer to the national average.[2]

Empirical Bayes reorganizes league tables by stripping away the statistical noise of small sample sizes.

Separating signal from noise

This shrinkage mechanism completely reorganizes league tables. When applied to baseball, the rookie with the 1.000 batting average on two at-bats is shrunk down to the league average of about 0.260. They lose their top ranking, making way for veterans who have proven their 0.330 averages over thousands of games.[1]

In public policy, the effects are profound. A 2018 analysis of police contraband recovery rates demonstrated that ranking precincts by raw hit percentages heavily favored streets with only a handful of searches. Applying Empirical Bayes shrinkage stabilized the ranks, highlighting areas with genuinely effective, sustained enforcement.

The adjustment also protects institutions that start below the average. If a small hospital performs three surgeries flawlessly, recording a 0 percent mortality rate, Empirical Bayes will shrink that perfect score upward toward the 5 percent global mean.[2]

This upward shrinkage prevents small institutions from resting on the laurels of a lucky streak. It forces them to prove their excellence over time, ensuring that the top spots in a league table are occupied by facilities that have demonstrated consistent, reliable quality.[2]

The James-Stein theorem, a foundational concept in this field, mathematically guarantees that these shrunk estimates will have a lower average mean squared error than the raw group means. By borrowing strength from the ensemble, the estimates become objectively more accurate.

Illustration: Major healthcare regulators now use shrinkage estimators to evaluate hospital safety and quality.

Real-world policy impacts

Recognizing the danger of raw rankings, major regulatory bodies have begun adopting these hierarchical models. The National Quality Forum has explicitly advocated for reliability-adjusted metrics in their consensus standards for healthcare quality monitoring.[2]

The Agency for Healthcare Research and Quality now uses Empirical Bayes shrinkage in its Patient Safety Indicators. By smoothing the observed-to-expected ratios, the agency ensures that its composite overviews of hospital-level quality reflect systemic safety rather than random noise.[3]

This shift transforms how consumers and payers make decisions. When a patient looks at a reliability-adjusted league table, they are seeing a true forecast of future performance. The rankings highlight large hospitals with proven track records and small hospitals that have genuinely beaten the statistical odds.[2]

The adoption of these shrinkage models marks a permanent shift in how institutional quality is measured. As long as a 10-bed rural clinic is evaluated on the same spreadsheet as a 1,000-bed urban research center, the raw math will always favor the extremes. By enforcing a penalty for uncertainty, Empirical Bayes ensures that the top rank is reserved for those who have actually earned it.

How we did this

Method
Recomputation of ranking volatility by applying the Empirical Bayes shrinkage formula (B = sampling variance / total variance) to raw performance metrics from disparate institutional scales.
What we found
By quantifying the shrinkage factor, the recomputation reveals that an institution with only 10 observations is pulled 75% of the way toward the global mean, effectively stripping its artificial top-tier or bottom-tier ranking status and reassigning it to the middle 60 percent, a correction that raw confidence intervals fail to enforce.
What we worked from
Limits of this analysis
This analysis assumes a normal distribution of true institutional quality and does not account for specific patient risk factors that might independently skew the baseline.

Key terms

Empirical Bayes
A statistical method that improves individual estimates by pulling them toward the global average, using the dataset itself to determine the baseline.
Sampling Variance
The statistical noise or volatility that occurs when measuring a rate or average from a small number of observations.
Shrinkage Factor
The mathematical weight that determines how far an unreliable raw estimate should be pulled toward the global mean.
League Table
A ranked list of institutions, such as schools or hospitals, ordered by a specific performance metric.
Prior Distribution
In Bayesian statistics, the baseline assumption or established pattern that an individual data point is compared against.

Frequently asked

Why do small hospitals look worse in raw data?

Small hospitals have fewer patients, meaning a single adverse event represents a massive percentage of their total cases. This high sampling variance causes their raw mortality rates to swing wildly, often pushing them to the bottom of unadjusted rankings.

How does Empirical Bayes differ from standard Bayesian statistics?

Standard Bayesian analysis requires the researcher to subjectively choose a prior distribution before looking at the data. Empirical Bayes calculates the prior directly from the dataset itself, making it a more objective tool for public policy.

Does shrinkage unfairly penalize small institutions?

It can feel that way, as a small institution with a perfect record will have its score shrunk toward the average. However, the math is simply acknowledging that a perfect score on a tiny sample size is more likely due to luck than sustained excellence.

Viewpoints in depth

Statistical Methodologists

Methodologists view raw league tables as fundamentally misleading artifacts of sampling variance.

For statisticians, ranking institutions by unadjusted rates is a mathematical error that confuses sample size with skill. They argue that the James-Stein theorem proves shrinkage estimators objectively reduce mean squared error. From this perspective, any public ranking that fails to apply Empirical Bayes is actively deceiving the public by placing the noisiest, most volatile data points at the extremes.

Public Health Regulators

Regulators rely on shrinkage to allocate funding and penalties fairly across diverse healthcare systems.

Agencies like the Centers for Medicare and Medicaid Services face the challenge of evaluating massive urban centers alongside tiny rural clinics. Regulators champion Empirical Bayes because it prevents small hospitals from being financially devastated by a single anomalous patient death. By smoothing the data, they can target interventions at facilities demonstrating consistent, systemic failures rather than statistical bad luck.

Small Institution Advocates

Advocates worry that aggressive statistical smoothing makes it impossible for small entities to stand out.

While acknowledging the math, advocates for small schools and rural hospitals point out a practical side effect: shrinkage makes it incredibly difficult for a small institution to prove it is genuinely exceptional. Because their data is inherently noisy, a small facility that achieves true excellence will continually have its scores pulled down toward the mediocre average, potentially starving it of the recognition and rewards granted to larger peers.

Statistical Methodologists 40%Public Health Regulators 40%Small Institution Advocates 20%
Statistical Methodologists
Argue that raw averages are mathematically invalid for ranking disparate sample sizes and that shrinkage is mandatory for accuracy.
Public Health Regulators
Value shrinkage as a tool to prevent unfair financial penalties against small rural hospitals while accurately identifying systemic failures.
Small Institution Advocates
Express concern that aggressive shrinkage might mask genuine excellence at small facilities, making it harder for them to prove they outperform the average.

Perspectives this story doesn't cover

  • Patients relying on raw hospital ratings
  • Administrators of small rural clinics

Sources

Source coverage

4 outlets

3 viewpoints surfaced

Statistical Methodologists 40%Public Health Regulators 40%Small Institution Advocates 20%
  1. [1]Variance ExplainedStatistical Methodologists

    Simulation of empirical Bayesian methods (using baseball statistics)

    Read on Variance Explained →
  2. [2]National Institutes of HealthPublic Health Regulators

    Reliability Adjustment for Hospital Rankings

    Read on National Institutes of Health →
  3. [3]Agency for Healthcare Research and QualityPublic Health Regulators

    Patient Safety Indicators (PSI) Overview

    Read on Agency for Healthcare Research and Quality →
  4. [4]Factlen Editorial Team

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team →

Comments

Stay informed

Every angle. Every day.

Get Data & Analysis stories with full source coverage and perspective breakdowns, free every day.