The Science of Superforecasting: Why the Brier Score Proves Prediction is a Trainable Skill
By decomposing predictions into calibration and resolution, researchers have demonstrated that forecasting is not an innate gift but a measurable, trainable process. The application of strictly proper scoring rules reveals that structured teams consistently outperform both individual experts and prediction markets.
- Behavioral Scientists
- Believe forecasting is a trainable cognitive skill.
- Market Theorists
- Argue that liquid prediction markets are the ultimate arbiters of probability.
- Algorithmic Aggregators
- Focus on mathematical models to extract certainty from diverse crowds.
Perspectives this story doesn't cover
- Institutional policymakers who rely on un-scored expert intuition
- Financial traders operating in high-liquidity prediction markets
At a glance
- The Brier score proves forecasting is a measurable, trainable skill rather than an innate talent.
- A strictly proper scoring rule forces forecasters to report honest probabilities by heavily penalizing overconfidence.
- Superforecasters consistently beat prediction markets by 15 to 30 percent in IARPA tournaments.
- The extremizing algorithm extracts hidden certainty from cognitively diverse teams, pushing aggregated forecasts closer to 0 or 100 percent.
- A single hour of cognitive-debiasing training can improve a forecaster's accuracy by 10 percent.
At the end of the second year of a sprawling US intelligence tournament, a volunteer named Doug recorded a Brier score of 0.14 across hundreds of geopolitical predictions. That number, a mathematical measure of probabilistic accuracy, made him the single best forecaster among 2,800 participants in the Good Judgment Project. It also proved a point that upends a century of conventional wisdom: accurate foresight is not an innate, mystical gift. It is a mechanical, trainable skill.[2]
The Good Judgment Project began in July 2011 as a response to a challenge from the Intelligence Advanced Research Projects Activity (IARPA). The agency wanted to know if it was possible to generate accurate, timely probabilistic forecasts for global events by aggregating the judgments of widely dispersed analysts. Led by researchers Philip Tetlock, Barbara Mellers, and Don Moore, the project recruited thousands of talented amateurs and subjected their predictions to rigorous statistical scoring.[2]
The foundation of this scoring is the Brier score, first proposed by Glenn W. Brier in 1950. The Brier score is a "strictly proper scoring rule" that measures the mean squared difference between a predicted probability and the actual binary outcome. If an analyst predicts a 70 percent chance of an event and it happens, the score is the square of 0.70 minus 1, or 0.09. If it does not happen, the score is the square of 0.70 minus 0, or 0.49.[1]
Because the Brier score is strictly proper, forecasters maximize their expected scores if, and only if, they state their true probabilities. There is no mathematical advantage to hedging or exaggerating. The penalty grows exponentially with overconfidence, meaning a forecaster who predicts 99 percent and is wrong is punished far more severely than one who predicts 60 percent and is wrong.[1]
To understand why some forecasters succeed, researchers decompose the Brier score into three distinct components: calibration, resolution, and uncertainty. Calibration, or reliability, measures how close the assigned probabilities are to the actual frequency of events. If a forecaster says 20 different events have a 10 percent chance of occurring, exactly two of them should happen.[1]
Resolution, meanwhile, measures the forecaster's ability to go out on a limb and make non-trivial predictions. If an event has a historical base rate of 50 percent, a forecaster who always guesses 50 percent will have perfect calibration but terrible resolution. A skilled forecaster must distinguish the 90 percent likelihoods from the 10 percent likelihoods. Uncertainty is simply the inherent randomness of the event itself, independent of the forecaster.[1]
Resolution, meanwhile, measures the forecaster's ability to go out on a limb and make non-trivial predictions.
Tetlock’s earlier research had famously concluded that the average political expert was "roughly as accurate as a dart-throwing chimpanzee." But the Good Judgment Project revealed a subset of people—dubbed "superforecasters"—who consistently beat the baseline. These individuals were not necessarily domain experts with classified clearance. They were simply highly numerate people who updated their beliefs frequently, broke intractable problems into smaller components, and maintained deep intellectual humility.[3]
The strongest counter-argument to the Good Judgment Project's team-based model is the prediction market. Economic theory dictates that a market with real liquidity—where participants buy and sell shares of future events—should perfectly aggregate all available information. If a market prices an event at 60 cents, the probability is 60 percent. Many economists argue that markets will always beat panels of experts.
Yet the data from the IARPA tournament showed something else. Teams of superforecasters consistently beat prediction markets by margins of 15 to 30 percent. They achieved this not just through individual brilliance, but through a mathematical aggregation technique known as the extremizing algorithm.[3]
The extremizing algorithm works by taking a team's aggregated probability estimate and pushing it closer to 0 percent or 100 percent. If three cognitively diverse forecasters independently conclude an event has a 70 percent chance of happening, the algorithm might push that aggregate to 85 percent. The reasoning is transparent: if people with entirely different biases, data sources, and life experiences all arrive at the same confident conclusion, they have collectively covered a massive amount of the hypothesis space.
This mathematical free lunch comes with a strict condition: it depends entirely on team diversity. If a team is a monolith where everyone shares the same background and reads the same reports, they should not be extremized at all. In a zero-diversity environment, the algorithm simply amplifies shared blind spots and systemic biases.
The implications of this research extend far beyond intelligence gathering. In a controlled experiment, the Good Judgment Project demonstrated that a single one-hour cognitive-debiasing training module—teaching probabilistic reasoning and the avoidance of hindsight bias—consistently improved forecasters' Brier scores by 10 percent. The barrier to accurate foresight is not a lack of innate talent, but a refusal to keep score.[2]
The science of forecasting proves that while the future remains uncertain, our ability to measure that uncertainty is a solved problem. The mathematical frameworks exist, and the cognitive training requires minimal time investment. The only remaining obstacle is whether institutions are willing to subject their own experts to the unforgiving, objective reality of a strictly proper scoring rule.[4]
Terms to know
- Brier Score
- A strictly proper scoring rule that measures the mean squared error of probabilistic forecasts.
- Calibration
- The degree to which a forecaster's assigned probabilities match the actual frequency of the events occurring.
- Resolution
- A forecaster's ability to confidently distinguish between highly likely and highly unlikely events, rather than clustering guesses near the base rate.
- Extremizing Algorithm
- A mathematical technique that pushes the aggregated probability of a diverse team closer to 0% or 100% to account for collective certainty.
- Strictly Proper Scoring Rule
- A scoring system where a forecaster can only maximize their expected score by reporting their genuinely honest probability estimate.
Sources
[1]WikipediaBehavioral ScientistsBrier score
Read on Wikipedia →
[2]WikipediaBehavioral ScientistsThe Good Judgment Project
Read on Wikipedia →
[3]80,000 HoursAlgorithmic AggregatorsPhilip Tetlock on how to predict the future
Read on 80,000 Hours →
[4]Factlen Editorial TeamBehavioral ScientistsSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
Comments
More in Opinion
See all →Information Theory
Why the Bekenstein Bound Proves That Exceeding a Physical Information Density Threshold Creates a Black Hole
8 sources
Market Concentration
The Minimum Efficient Scale: Why Economies of Scale Guarantee That Only a Few Giant Firms Can Survive in an Industry
9 sources
Algorithm Theory
The Mathematical Impossibility of a Universal Optimizer: Why the No Free Lunch Theorem Means No Algorithm Is Inherently Superior
5 sources
North American Trade
By Refusing to Renew USMCA, Did the US Just Weaponize Trade Uncertainty as a Permanent Industrial Policy Tool?
7 sources
Every angle. Every day.
Get Opinion stories with full source coverage and perspective breakdowns delivered to your inbox.




