Skip to main content
ExplainerA/B TestingMicrosoft· 7 min read· in Data & Analysis

Sample Ratio Mismatches Invalidate A/B Test Results Through Uncorrectable Selection Bias

When an online experiment fails to deliver its intended traffic split, the discrepancy often masks differential user attrition rather than random chance. This sample ratio mismatch introduces a severe selection bias that mathematical reweighting cannot fix, forcing data science teams to discard the results entirely.

By Logan Price

In short

  • Sample ratio mismatch occurs when the observed traffic split in an A/B test diverges significantly from the intended allocation, violating the core assumption of randomization.
  • The imbalance is rarely a random error; it typically indicates differential user attrition where one variant loses users to bugs, latency, or tracking failures.
  • Because the missing users are not a random sample, their absence introduces a severe selection bias that cannot be corrected by mathematically reweighting the surviving data.

The foundation of online experimentation rests on a simple coin flip. When a technology company tests a new feature, it randomly assigns users to a control group or a treatment group, expecting the law of large numbers to make both populations identical. That initial randomization is the only mechanism that allows analysts to attribute any subsequent behavioral changes directly to the new feature.[5]

But the coin does not always land fairly. When the final tally of users in an A/B test diverges significantly from the intended allocation—such as a planned 50/50 split returning 50.2 percent in one arm and 49.8 percent in the other—the experiment suffers from a sample ratio mismatch.[1][2]

While a two-tenths of a percent deviation sounds trivial, at the scale of modern software, it represents a catastrophic failure. Across a sample of 1.64 million users, a 50.2 to 49.8 percent split is a statistical impossibility under fair randomization.[2]

"Sample ratio mismatch is evidence of selection bias," notes the experimentation platform Statsig in its 2023 technical documentation. If the assignment is not truly random, other hidden variables inevitably separate the two groups, destroying the causal link between the product change and the observed outcome.[5]

The phenomenon is remarkably common across the software industry. In a landmark 2019 paper presented at the Knowledge Discovery and Data Mining conference, researchers from Microsoft Corporation and Booking.com revealed that approximately 6 percent of all A/B tests conducted on their platforms exhibited a sample ratio mismatch.[3]

Approximately 6 percent of all A/B tests exhibit a sample ratio mismatch.

"I can recall hundreds of SRMs. We consider it one of the most severe data quality issues we can detect," one Microsoft analyst reported in the 2019 study. The researchers classified the mismatch not as a minor technical glitch, but as a fatal symptom of underlying data corruption that renders the resulting analysis completely untrustworthy.[3]

The Mechanism of Differential Attrition

To understand why a slight numerical imbalance destroys an experiment, analysts look to the mechanism that causes it. A sample ratio mismatch rarely occurs because the random number generator itself failed. Instead, it happens because users drop out of the experiment pipeline at different rates after the assignment occurs but before the telemetry logs their presence.[1][6]

Consider a software update that inadvertently adds 500 milliseconds of latency to a mobile application's loading screen. The randomization engine assigns users perfectly evenly at the server level, but users in the treatment group experience a noticeably slower application on their devices.[1]

Frustrated by the delay, a fraction of those treatment users close the application before the tracking pixel fires to record their participation. The control group, experiencing normal load times, logs its users reliably. The resulting data shows fewer users in the treatment arm, creating a classic sample ratio mismatch.[1][6]

This differential attrition means the users who vanished are not a random cross-section of the audience. They are specifically the users most sensitive to performance degradation. Because their data is missing entirely, the surviving treatment group looks artificially resilient, skewing the final conversion metrics upward.[6]

A seemingly trivial 0.2 percent deviation represents a massive statistical anomaly at scale.

The Survivorship Bias Analogy

Microsoft researchers compare the phenomenon to the famous survivorship bias identified by statistician Abraham Wald during World War II. When the military asked Wald to optimize aircraft armor based on the bullet holes found on returning bombers, he famously recommended reinforcing the areas with no damage.[3]

Wald realized that the planes taking hits in critical areas simply never returned to be measured. In online experimentation, the users who encounter a broken feature, a redirect loop, or a bot-filtering error are the missing planes.[3]

"Just as Abraham Wald could not conduct a complete analysis of aircraft survival without considering those planes which did not return, so too do A/B testers need to be aware of missing users in their experiments," the Microsoft research team explained in their 2020 methodology review.[3]

When an experiment loses users disproportionately from one arm, the remaining population is fundamentally altered. Analyzing the survivors without accounting for the missing cohort guarantees a biased conclusion, often making a harmful product change appear highly successful.[3][7]

Why Reweighting Fails

When faced with a sample ratio mismatch, product teams often attempt to rescue the experiment by mathematically reweighting the data to force a 50/50 split. This approach assumes that the missing users would have behaved exactly like the surviving users, which is precisely the assumption that differential attrition violates.[2][4]

Data scientist Lukas Vermeer demonstrates the mathematical impossibility of correcting this bias through extreme-value bounding. If a control group logs 10,000 users while the treatment group logs only 9,000, the experiment is missing 1,000 counterfactual treatment users.[4]

To calculate the true bounds of the treatment effect, analysts must impute the best-case and worst-case scenarios for those 1,000 missing individuals. If the metric is a conversion rate, the worst-case scenario assumes zero conversions among the missing cohort, while the best-case scenario assumes a 100 percent conversion rate.[4][7]

Extreme-value bounding demonstrates how missing counterfactuals stretch the confidence interval across zero.

Because the missing population is large enough to trigger the mismatch alert, imputing these extreme values invariably stretches the confidence interval across zero. The true treatment effect could be entirely negative or highly positive, proving mathematically that the point estimate is meaningless and cannot be salvaged.[4][7]

Detecting the Mismatch

Because eyeballing percentages is unreliable, experimentation platforms rely on strict statistical tests to flag imbalances. The industry standard is the Pearson's chi-square goodness-of-fit test, which calculates the probability that the observed distribution occurred by chance under a fair allocation.[1][5]

Platforms deliberately set the sensitivity threshold for this check much stricter than the standard 0.05 p-value used for reading experiment results. A 5 percent threshold would raise a false alarm on one in every twenty experiments, causing teams to ignore the warning entirely.[1]

Instead, systems like Microsoft's experimentation platform use a highly conservative threshold of p < 0.0005, while others standardize around p < 0.001. When an experiment trips an alarm at this level, the imbalance is almost certainly a genuine systemic fault rather than bad luck.[1][2]

The scale of the experiment heavily influences this calculation. In a small test with 1,000 users per arm, a 53 to 47 percent split might not trigger the chi-square threshold. However, at 1,000,000 users, a deviation of just 0.098 percent is statistically overwhelming and demands immediate investigation.[1][2]

The statistical significance of a traffic imbalance scales exponentially with the total number of users.

Diagnosing the Root Cause

Once an experimentation platform flags a sample ratio mismatch, the analysis must halt. "Sample ratio mismatch is a stop sign for causal interpretation, not a cosmetic warning beside the result," experimentation consultant Atticus Li noted in a 2026 methodology brief.[2]

The diagnostic process requires tracing the user journey backward to find the exact point where the ratio breaks. Analysts must examine the assignment logs, the execution triggers, the telemetry pipelines, and the data filtering rules to isolate the leak.[2][6]

Common culprits include targeting rules that inadvertently exclude specific user segments from one variant, client-side execution failures caused by ad blockers, or bot-detection algorithms that aggressively filter out highly engaged users in a successful treatment arm.[3][6]

In one documented case at Microsoft, a treatment variant increased user engagement so dramatically that the platform's bot-detection system classified the genuine users as automated traffic and filtered them out. The resulting sample ratio mismatch initially made the highly successful feature look like a failure.[3]

The Financial Stakes

The financial stakes of ignoring these warnings are massive. In a case study known as the "$10 Million Mirage," a company ran an experiment that appeared to generate a 2 percent uplift in weekly revenue, projecting an annualized gain of $10 million.[8]

However, the experiment carried a subtle sample ratio mismatch, showing a 49.95 to 50.05 percent split. Upon investigation, the team discovered that internal employees had been inadvertently assigned exclusively to the treatment group.[8]

Because employees engaged with the product far more frequently than standard users, their presence artificially inflated the revenue metrics. Once the employee data was removed and the mismatch resolved, the actual incremental revenue dropped to zero, proving that a 1 percent imbalance can easily fabricate a 2 percent financial gain.[8]

A 1 percent traffic imbalance can easily fabricate a 2 percent financial gain.

Until the root cause is identified and the system is fixed, the data remains compromised. The only scientifically defensible response to an uncorrectable sample ratio mismatch is to discard the results, resolve the underlying instrumentation error, and launch a completely new experiment.[2][7]

How we did this

Method
Extreme-value bounding of missing counterfactuals to quantify uncorrectable bias
What we found
By imputing both a 0% and 100% conversion rate for the 1,000 missing treatment users, the bounding exercise proves that the true treatment effect spans across zero regardless of the surviving users' behavior, demonstrating mathematically why post-hoc reweighting cannot salvage an experiment with differential attrition.
What we worked from
Limits of this analysis
This bounding technique only proves that the point estimate is invalid; it cannot reconstruct what the true treatment effect would have been had the attrition not occurred.

Definitions

Sample Ratio Mismatch (SRM)
A statistically significant difference between the expected and actual traffic allocation in a controlled experiment.
Selection Bias
A statistical error that occurs when the participants included in an analysis are not truly representative of the intended population.
Counterfactual
The unobserved outcome of what would have happened to a specific user had they been assigned to a different experimental condition.
Extreme-Value Bounding
A diagnostic technique that calculates the best-case and worst-case scenarios for missing data to determine if an effect remains statistically significant.

Questions & answers

Can I fix a sample ratio mismatch by randomly dropping users from the larger group?

No. Dropping users from the larger group only masks the symptom without addressing the underlying selection bias. The remaining users in the smaller group are still a skewed population, meaning the comparison remains invalid.

Why is the p-value threshold for SRM so much stricter than for the main experiment?

Because the SRM check runs on every single experiment automatically, using a standard 0.05 threshold would trigger a false alarm on 5 percent of all valid tests. A stricter threshold like 0.001 ensures that alarms only sound for genuine systemic faults.

Does a sample ratio mismatch always mean the new feature is broken?

Not necessarily. While bugs and latency are common causes, an SRM can also occur if a highly successful feature causes users to engage so rapidly that they trigger bot-detection filters, inadvertently excluding them from the analysis.

Analysis by camp

Experimentation Platform Engineers

Focus on building robust telemetry and strict automated thresholds to prevent corrupted data from reaching decision-makers.

Engineers who design experimentation platforms view sample ratio mismatch as a systemic infrastructure failure rather than a mere statistical anomaly. They advocate for building automated, end-to-end diagnostic pipelines that halt the analysis of any test failing a strict chi-square check. For this camp, the priority is ensuring that telemetry, assignment logs, and bot-filtering algorithms remain perfectly synchronized, preventing flawed data from ever informing a product decision.

Statistical Methodologists

Emphasize the mathematical impossibility of reweighting biased samples and advocate for bounding techniques to prove invalidity.

Methodologists approach sample ratio mismatch through the lens of causal inference and selection bias. They strongly oppose any attempts to 'salvage' a mismatched experiment through post-hoc reweighting or propensity matching, arguing that differential attrition fundamentally destroys the randomized counterfactual. Instead, they rely on extreme-value bounding to mathematically demonstrate that imputing outcomes for the missing users stretches the confidence interval across zero, proving that the point estimate is entirely meaningless.

Experimentation Platform Engineers 50%Statistical Methodologists 50%
Experimentation Platform Engineers
Focus on building robust telemetry and strict automated thresholds to prevent corrupted data from reaching decision-makers.
Statistical Methodologists
Emphasize the mathematical impossibility of reweighting biased samples and advocate for bounding techniques to prove invalidity.

Perspectives this story doesn't cover

  • Product Managers relying on flawed data
  • End Users experiencing broken variants

Sources

Source coverage

8 outlets

2 viewpoints surfaced

Experimentation Platform Engineers 50%Statistical Methodologists 50%
  1. [1]Bell StatisticsStatistical Methodologists

    What is sample ratio mismatch (SRM)?

    Read on Bell Statistics →
  2. [2]Atticus LiStatistical Methodologists

    Diagnosing Sample Ratio Mismatch

    Read on Atticus Li →
  3. [3]Microsoft ResearchExperimentation Platform Engineers

    Diagnosing Sample Ratio Mismatch in Online Controlled Experiments

    Read on Microsoft Research →
  4. [4]Lukas VermeerStatistical Methodologists

    What is Sample Ratio Mismatch?

    Read on Lukas Vermeer →
  5. [5]StatsigExperimentation Platform Engineers

    Sample ratio mismatch is evidence of selection bias

    Read on Statsig →
  6. [6]KameleoonExperimentation Platform Engineers

    Sample Ratio Mismatch (SRM) defined

    Read on Kameleoon →
  7. [7]Factlen Editorial TeamStatistical Methodologists

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team →
  8. [8]Towards Data ScienceStatistical Methodologists

    The $10 Million Mirage

    Read on Towards Data Science →

Comments

Stay informed

Every angle. Every day.

Get Data & Analysis stories with full source coverage and perspective breakdowns, free every day.