Data Normalization in Choropleth Maps: Why Using Counts Instead of Rates Distorts Geographic Trends
Choropleth maps use color to visualize geographic data, but mapping raw counts instead of normalized rates creates misleading population maps. Adjusting for population or area is mathematically required to reveal true spatial patterns.
By Logan Price
- Cartographic Standards
- Advocates for strict adherence to data normalization and the use of graduated symbols for raw counts.
- Statistical Methodology
- Focuses on the mathematical accuracy of spatial data and the prevention of ecological fallacies.
- Open Knowledge & Synthesis
- Aggregates historical context and broad consensus on the limitations of thematic mapping.
Perspectives this story doesn't cover
- Automated Software Developers
- Cognitive Psychologists
Choropleth maps color-code geographic areas to show data trends, but when they use raw counts instead of normalized rates, they simply recreate population maps. Normalizing data—dividing counts by an underlying population or area—is the only mathematical way to reveal true geographic patterns. Without this adjustment, a map showing the total number of hospitals, crimes, or disease cases will always highlight the most populous regions, misleading the viewer into seeing a crisis where there is only a crowd.[3][4]
The choropleth map has been a foundational tool for geographic data visualization since 1826, when Baron Pierre Charles Dupin created the first tinted map to depict educational availability across French departments. Today, these maps are ubiquitous in news reports, election coverage, and public health dashboards. They operate by filling established geographic polygons—such as counties, states, or census tracts—with a sequential color scale based on a specific variable.[5]
However, the visual power of these maps is easily compromised by a failure to normalize the underlying data. "If you forget to normalize data for a choropleth map... you often end up showing population centers, instead of the phenomenon that you're trying to measure," explain the authors of Hands-On Data Visualization. A map of raw unemployment figures, for instance, will invariably shade California and Texas in the darkest colors simply because they contain the most workers, obscuring the actual rate of joblessness in smaller states.[3][4]
The mathematical distortion introduced by raw counts is severe. Consider historical COVID-19 data from late 2020: the United States recorded 9.61 million cases against an estimated population of 328.2 million, while Belgium recorded 0.49 million cases against a population of 11.5 million. On a raw-count choropleth map, the United States appears to have an infection burden nearly 20 times worse than Belgium, dominating the visual hierarchy.[3]
Normalizing that same data to a per-capita rate reverses the narrative entirely. When adjusted to cases per 100,000 residents, the United States stood at 2,928, while Belgium stood at 4,260. The raw count approach overstates the relative US incidence by a factor of 28.8 compared to the normalized rate. This derivation demonstrates that failing to adjust for population does not merely blur geographic insights; it can mathematically invert the true distribution of a phenomenon.[3][6]
Normalizing that same data to a per-capita rate reverses the narrative entirely.
To correct this, cartographers employ three primary methods of normalization. The most common is population normalization, which divides a raw count by the total number of people in the polygon to generate a per-capita rate. This method is essential for demographic data, public health statistics, and crime rates, ensuring that a city of 5 million people is judged on the same scale as a town of 50,000.[2][4]
The second method is area normalization, which divides a total count by the physical size of the geographic unit to calculate density. This is particularly useful when mapping environmental phenomena, agricultural yields, or the distribution of physical infrastructure. "A properly normalized map will show variables independent of the size of the polygons," notes the encyclopedic consensus on spatial analysis, preventing vast but empty regions from dominating the visual field.[5]
The third approach involves calculating proportions or percentages, where a subgroup is divided by a grand total. Mapping the percentage of wealthy households out of all households, or the share of the electorate that voted for a specific candidate, provides a standardized metric that remains consistent regardless of the absolute numbers involved.[5]
Even with perfectly normalized data, choropleth maps suffer from an inherent visual bias: region size. Large geographic areas command disproportionate attention from the human eye, even if their data values are unremarkable. This can obscure critical patterns in densely populated, geographically small urban centers. Because choropleths aggregate data to the boundaries of the polygon, they also mask any internal variation within that region, presenting a uniform color that implies a uniform reality.
The consequences of ignoring these cartographic rules are not merely academic. During the height of the COVID-19 pandemic, one study found that more than 50 percent of the dashboards hosted by state governments in the United States failed to employ normalization in their choropleth maps. This widespread methodological error contributed to public confusion and misallocated resources, as raw-count maps directed attention toward dense cities rather than areas with genuinely high transmission rates.[5]
When raw counts must be displayed, spatial analysts recommend abandoning the choropleth format entirely. "The general rule to follow is that raw values should be mapped with graduated or proportional symbols, and derived values should be mapped with color," advises Esri's cartography documentation. Graduated symbol maps place a scalable circle over each region, allowing the size of the symbol to represent the magnitude of the count without letting the polygon's physical area distort the perception of the data.[2]
The integrity of geographic data visualization relies entirely on the mathematical foundation beneath the colors. As mapping tools become more accessible, the responsibility shifts to the creator to ensure the data is properly scaled. Until software defaults universally enforce normalization, readers must remain skeptical of any shaded map that fails to specify whether it is showing a rate, a ratio, or simply a count of the people living there.[1][4]
- 28.8x
- Distortion factor of raw counts
- 9.61 million
- US raw cases (late 2020)
- 4,260
- Belgium cases per 100k
- >50%
- US state COVID dashboards lacking normalization
- 1826
- Year of first choropleth map
Limits of the evidence
- How frequently modern automated mapping software still defaults to raw counts without warning the user.
- The exact degree to which region-size bias alters public perception of normalized data compared to alternative formats like cartograms.
- Whether unclassed choropleth maps (which use a continuous color gradient) reduce misinterpretation compared to classed maps with rigid data buckets.
Sources
[1]ONS Service ManualStatistical MethodologyOther charts and maps: Choropleth maps – Data visualisation
Read on ONS Service Manual →
[2]EsriCartographic StandardsNormalization for choropleth maps
Read on Esri →
[3]Hands-On Data VisualizationStatistical MethodologyNormalize Choropleth Map Data
Read on Hands-On Data Visualization →
[4]FeltCartographic StandardsChoropleth maps: Color-coding patterns without misleading your audience
Read on Felt →
[5]WikipediaOpen Knowledge & SynthesisChoropleth map
Read on Wikipedia →
[6]Factlen Editorial TeamOpen Knowledge & SynthesisSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
Comments
More in Data & Analysis
See all →Chart Geometry
The Geometry of Deception: Why Bar Charts Require a Zero Baseline While Line Charts Do Not
7 sources
Evaluation Metrics
How the Quadratic Penalty in RMSE Forecast Evaluation Punishes Outliers Compared to MAE's Linear Loss
5 sources
Survey Methodology
Why Complex Survey Designs Lose Statistical Power: Inside the Design Effect Penalty
9 sources
Search Algorithms
BM25 vs. Dense Retrieval: The Accuracy and Latency Trade-offs in Search Ranking
2 sources
Every angle. Every day.
Get Data & Analysis stories with full source coverage and perspective breakdowns delivered to your inbox.




