Evidence Pack: How Multilevel Regression and Poststratification (MRP) Estimates Local Opinion from National Polls
By combining Bayesian regression with rigid Census benchmarks, MRP allows statisticians to extract highly accurate, localized estimates from skewed and non-representative survey samples.
By Ishani Patel
- Survey Methodologists
- Value MRP for its ability to reduce variance and salvage non-representative samples.
- Commercial Pollsters
- Balance the statistical benefits of MRP against its high computational and labor costs.
- Election Forecasters
- Utilize MRP for localized seat projections but warn about the risks of flawed Census frames.
Perspectives this story doesn't cover
- Traditional Design-Based Survey Statisticians
- Census Bureau Demographers
Key points
- MRP is a two-step statistical method used to extract accurate, localized estimates from non-representative survey samples.
- The regression step uses partial pooling to estimate opinions for specific demographic groups without the high variance of small sample sizes.
- The poststratification step weights those predictions against rigid population benchmarks, such as Census data.
- While national-level accuracy is similar to traditional weighting, MRP drastically reduces the margin of error for small subgroups.
- The technique is computationally expensive, requiring a separate regression model for every individual question on a survey.
- ±4 points
- MRP margin of error for Hispanic voters (vs ±7 points for raking)
- 93%
- UK constituencies correctly predicted by YouGov's 2017 MRP model
- < 5%
- Modern telephone survey response rates, driving the need for MRP
In 2012, when the Pew Research Center attempted to estimate the voting behavior of Hispanic Americans using a standard opt-in online poll, the traditional weighting method known as raking produced a margin of error of plus or minus 7 percentage points. That wide interval made it difficult to draw precise conclusions about the subgroup. However, when statisticians ran the exact same raw survey data through a technique called Multilevel Regression and Poststratification (MRP), the margin of error shrank to just plus or minus 4 points. The national-level accuracy remained identical at plus or minus 2 points, but the subgroup precision nearly doubled.[1]
This mathematical leverage is why MRP—often affectionately called "Mister P" by statisticians—has become the gold standard for modern public opinion research. As response rates for traditional random-digit dialing have plummeted below 5 percent, pollsters have increasingly relied on large, non-representative online panels. Extracting a meaningful signal from these skewed samples requires aggressive statistical correction.
"Traditional polling methods often struggled with accurately representing small or niche groups within larger populations," notes the Australian research firm DemosAU. This limitation forced researchers to either ignore local trends or spend millions fielding massive, localized surveys. MRP bypasses this barrier by mathematically decoupling the survey sample from the population structure.
The methodology solves the non-representative sample problem in two distinct steps, beginning with multilevel regression. Instead of simply averaging the responses of the people who happened to take the survey, the model treats the data as a hierarchical structure. It estimates how various demographic traits—such as age, gender, education, and geographic location—interact to predict an individual's response.[2]
Crucially, this regression step employs a Bayesian technique known as "partial pooling." If a survey only captures a handful of young Hispanic voters in Ohio, traditional cross-tabulations would yield a highly volatile estimate for that specific group. MRP, however, allows that sparse cell to borrow statistical strength from young voters nationally, Hispanic voters in neighboring states, and Ohioans overall.[3]
This smoothing prevents small sample sizes from producing wild, noisy outliers. As the Pew Research Center explains, "The main advantage of MRP is that it lets researchers perform these kinds of complex adjustments without the same penalty to precision." The model learns the underlying relationships between demographics and opinions, rather than relying on raw headcounts.[1]
This smoothing prevents small sample sizes from producing wild, noisy outliers.
The second step is poststratification. Once the regression model has predicted the likely opinion of every possible demographic combination—often creating thousands of distinct "cells"—the algorithm weights each cell by its actual frequency in the true target population.
If official Census data shows that 35-year-old college-educated women make up exactly 1.2 percent of a state's population, their predicted opinion is weighted at exactly 1.2 percent of the final state estimate. This rigid recalibration occurs regardless of whether that demographic made up 5 percent of the survey sample or 0.1 percent.
The result is a synthetic but highly accurate projection of public opinion. By mapping modeled preferences onto rigid demographic benchmarks, MRP allows researchers to generate reliable state-by-state or district-by-district estimates from a single national survey.[3]
The technique gained mainstream fame in the United Kingdom when the polling firm YouGov used it to correctly predict the outcome in 93 percent of parliamentary constituencies during the 2017 general election. In the United States, researchers use MRP to estimate local health outcomes, combining survey data with demographic benchmarks to map disease prevalence at the census-tract level.[2][3]
However, the methodology comes with strict mathematical trade-offs. The most significant limitation is computational cost. Because MRP builds a predictive model for the outcome itself, researchers must run a separate, computationally intensive regression for every single question on a survey.[1]
Traditional raking, by contrast, calculates a single weight for each respondent that can be applied across a 40-question survey instantly. "If you have a survey with 40 questions, you would need to run MRP 40 times," the Pew Research Center noted, citing practicality as the primary barrier to universal adoption.[1]
Furthermore, MRP is entirely dependent on the quality of its poststratification frame. If the underlying Census data is outdated, or if the model omits a variable that strongly drives opinion—such as political party registration—the poststratification step will confidently project a biased result. The model can only correct for the demographic skews it is explicitly told to measure.[3]
Despite these constraints, the variance reduction MRP offers is transforming how polling organizations handle sparse data. By mathematically rebuilding the population from the ground up, statisticians at firms like YouGov and Pew can now extract high-resolution, localized insights from the noisy, fragmented reality of modern polling.[4]
How we got here
1997
Andrew Gelman and Thomas Little publish foundational research on poststratification using hierarchical logistic regression.
2012
Statisticians successfully use MRP on highly skewed Xbox user data to accurately predict the US presidential election.
2017
YouGov brings MRP into the mainstream by correctly predicting 93 percent of UK parliamentary constituencies.
2018
Pew Research Center publishes a comprehensive comparison showing MRP halves the margin of error for demographic subgroups compared to traditional raking.
What we don’t know
- How the increasing refusal rates across all demographics will eventually degrade the baseline survey data beyond what multilevel regression can mathematically salvage.
- Whether the computational costs of running separate Bayesian models for every survey question will drop enough to make MRP the default for daily commercial polling.
Sources
[1]Pew Research CenterSurvey MethodologistsComparing MRP to raking for online opt-in polls
Read on Pew Research Center →
[2]National Institutes of HealthSurvey MethodologistsMultilevel Regression and Poststratification: A Modeling Approach to Estimating Population Quantities From Highly Selected Survey Samples
Read on National Institutes of Health →
[3]WikipediaElection ForecastersMultilevel regression with poststratification
Read on Wikipedia →
[4]Factlen Editorial TeamSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
Comments
More in Data & Analysis
See all →Feature Selection
How the L1 Penalty in Lasso Regression Forces Coefficients to Zero for Feature Selection
6 sources
Constitutional Law
The U.S. Constitution Is the Second-Hardest to Amend in the Democratic World
3 sources
Statistical Methods
How the Bootstrap Method Uses Resampling to Estimate the Sampling Distribution of Any Statistic
6 sources
Global Wealth Gap
The Evidence Pack: World Bank Data Shows Half of Developing Economies Failing to Narrow Income Gap Since 2019
5 sources
Every angle. Every day.
Get Data & Analysis stories with full source coverage and perspective breakdowns delivered to your inbox.




