Skip to main content
ExplainerLoss FunctionsRetail Analytics· 7 min read· in Data & Analysis

Minimizing Absolute Error Targets the Median Rather Than the Mean: Why MAE Evaluation Causes Systematic Underforecasting in Right-Skewed Data

Mean Absolute Error (MAE) is widely used to evaluate forecasting models because it treats all errors equally, but mathematically, minimizing MAE forces a model to predict the median. In right-skewed datasets—such as retail demand, revenue, or time-on-site—the median is strictly lower than the mean, causing MAE-optimized models to systematically underforecast aggregate totals.

By Ishani Patel

In short

  • Mean Absolute Error applies a linear penalty to errors, meaning the mathematical target that minimizes it is the median.
  • In right-skewed datasets like retail demand, the median is strictly lower than the mean.
  • Aggregating MAE-optimized item-level forecasts results in a massive structural underforecast of total volume.

When data scientists and demand planners build a new forecasting model, one of their first decisions is choosing the metric that will grade its performance. They often select Mean Absolute Error (MAE) because it is highly intuitive to business stakeholders. An MAE of 10 simply means the model is off by 10 units on average.[1]

However, this seemingly innocuous choice embeds a hidden mathematical assumption that dictates how the model will behave in production. Unlike Mean Squared Error (MSE), which penalizes large errors exponentially, MAE applies a strict linear penalty to every miss. Because it treats all errors equally, the mathematical target that perfectly minimizes MAE is the median of the distribution.[3][5]

If a forecaster optimizes a machine learning model strictly for MAE, they are explicitly training it to find the median rather than the mean. This distinction is often overlooked during the model training phase, where teams assume that minimizing any error metric will naturally lead the model toward the true expected value.[1][2]

The divergence between the mean and the median only becomes a critical failure point because real-world business data is rarely symmetrical. Retail demand, website traffic, and revenue figures are typically right-skewed. They feature many small values clustered near zero, accompanied by a long tail of occasional massive spikes.[1]

In right-skewed data, the median is strictly lower than the mean, creating a structural gap.

In any right-skewed distribution, the median sits strictly lower than the mean. Because the data is pulled upward by the long tail of extreme values, the mathematical average is dragged to the right, while the median remains anchored near the bulk of the smaller observations.[1][6]

When a model is trained to predict the median on right-skewed data, it introduces a structural negative bias relative to the mean. On a per-item or per-day basis, the forecast looks highly accurate, often beating MSE-trained models in backtesting. The problem only becomes visible when those individual forecasts are aggregated.[1]

The Aggregation Trap

The fundamental issue is that while means are additive, medians are not. If a system forecasts the mean demand for every item in a catalog, summing those individual forecasts will yield the correct total demand for the entire warehouse.[1][4]

The median does not behave this way. Because the median is lower than the mean in skewed data, every individual MAE-optimized forecast is slightly too low compared to the expected average demand. When tens of thousands of these item-level forecasts are summed together, the small negative biases do not cancel out.[1]

Instead, they accumulate into a massive underforecast of total volume. A supply chain relying on these aggregated MAE forecasts will systematically under-order stock, leading to widespread shortages despite the model reporting excellent accuracy scores at the lowest level of granularity.[1]

While MAE forecasts perform well individually, they systematically underforecast when aggregated.

This phenomenon is particularly severe in intermittent demand forecasting, such as spare parts, luxury goods, or highly seasonal items. These datasets are often zero-inflated, meaning that a product might sell zero units on the vast majority of days, punctuated by sudden, unpredictable purchases.[2]

If an item sells zero units on more than half of the days in a dataset, the true mathematical median of that demand distribution is exactly zero. A model minimizing MAE will quickly learn that predicting zero every single day yields the lowest possible error score.[2]

While a flat zero forecast is mathematically optimal under an MAE evaluation, it is an operational disaster. Saying that a warehouse will not sell anything in the foreseeable future is a safe statistical strategy for intermittent data, but it guarantees that the business will never hold the inventory required to capture actual sales.[2]

The Mathematics of the Penalty

To understand why MAE forces this behavior, forecasters must examine the penalty surface of the metric. When evaluating a probabilistic forecast, the absolute error is minimized exactly at the point where the probability mass is split evenly in half.[3]

If a model predicts a value higher than the median, it increases the error for all the actual observations that fall below that point. Shifting the estimate up or down by one unit directly increases or decreases the resulting expected absolute error based purely on the count of observations on either side.[3]

Mean Squared Error operates on an entirely different principle. Because it squares the differences between the prediction and the actual outcome, MSE heavily penalizes large misses. To minimize this quadratic penalty, the model must shift its prediction toward the long tail of extreme values, landing exactly on the mean.[1][5]

Different loss functions optimize for entirely different mathematical targets.

This makes MSE the correct choice when the financial impact of a forecasting error scales disproportionately, or when aggregate totals matter more than item-level precision. By training models on MSE instead of MAE, baseline predictions on the item level are slightly higher, preventing the systematic underforecasting that appears after aggregation.[1]

Balancing Robustness and Accuracy

However, switching entirely to MSE introduces its own set of complications. Because MSE squares every error, it is highly sensitive to extreme outliers. A single corrupted data point or a massive, unrepeatable anomaly will aggressively distort an MSE-trained model as it attempts to minimize the quadratic penalty.[5]

Machine learning engineers often favor MAE precisely because of its outlier resistance. In noisy datasets, MAE provides a stable baseline that ignores extreme anomalies, ensuring that the model's core logic is not ruined by a handful of bad records.[5][6]

To resolve the tension between MAE's stability and MSE's unbiased aggregation, engineering teams often turn to hybrid loss functions. The Huber loss function is a popular alternative, acting like MAE for large errors to resist outliers, while behaving like MSE for small errors to maintain a mean-centric target.[5]

Another advanced approach involves using the Tweedie deviance family of cost functions. Tweedie regression is specifically designed for long-tailed, positive-only data, such as insurance claims, rainfall events, or retail sales, making it a natural fit for demand forecasting where negative values are impossible.[4]

Optimizing for Tweedie deviance allows forecasters to tune the penalty surface, assuming that the distribution is Tweedie-distributed around the point estimate rather than strictly normal. This provides a mathematical middle ground that accounts for the right-skew without defaulting to a pure median target.[4]

MSE applies a quadratic penalty to large errors, forcing the model to chase the long tail.

Teams can also employ asymmetric loss functions, such as quantile loss or pinball loss. These metrics allow forecasters to explicitly penalize under-prediction more heavily than over-prediction, which aligns well with retail environments where the business cost of a stockout exceeds the cost of holding excess inventory.[2][4]

The Role of Bias Correction

When teams are locked into using MAE or its percentage-based cousin, MAPE, due to legacy systems or stakeholder preferences, post-hoc bias correction becomes necessary. This involves training the model on the median-targeting metric and then mathematically adjusting the outputs upward to recover the mean.[4]

One common correction method is sales weighting, where the training data is weighted by the actual sales volume. By applying a square-root or linear sales weight to the observations, the model is forced to care more about high-volume days, dragging the median prediction closer to the true mean.[4]

While sales weighting is easy to implement in most machine learning libraries, it requires extensive experimentation to find the optimal balance. If the weighting is too aggressive, the negative bias is overcorrected, leading to severe over-forecasting and deteriorating the overall accuracy scores.[4]

Ultimately, the choice of evaluation metric is not just a technical detail; it is a business decision that shapes inventory, staffing, and revenue projections. A metric that looks perfect on a dashboard can silently erode operational efficiency if its mathematical properties do not align with the company's goals.[1]

Illustration: Aggregate forecasting accuracy dictates inventory levels across entire supply chain networks.

Data science teams must move beyond default metrics and explicitly define whether their downstream systems require the mean or the median. If the forecasts will be aggregated to drive high-level planning, targeting the mean is a structural necessity, regardless of how intuitive MAE might appear.[1][2]

As machine learning forecasting matures, the gap between how a model is trained and how it is evaluated must be closed. By understanding the deep link between MAE, the median, and right-skewed data, organizations can stop chasing misleading accuracy scores and start generating forecasts that actually add up.[6]

How we did this

Method
Recomputation of optimal constant forecasts on a synthetic right-skewed log-normal distribution to compare the penalty surfaces of MAE versus MSE.
What we found
A model perfectly minimizing MAE on a highly right-skewed dataset systematically underforecasts the aggregate sum, a structural shortfall that disappears entirely when minimizing MSE.
What we worked from
Limits of this analysis
Assumes a purely log-normal distribution; real-world data may have different skew profiles or bounded extremes.

Jargon, explained

Mean Absolute Error (MAE)
A metric that measures the average absolute difference between predicted and actual values, treating all errors equally regardless of direction.
Mean Squared Error (MSE)
A metric that squares the differences between predicted and actual values before averaging them, heavily penalizing large errors.
Right-Skewed Distribution
A dataset where most values are clustered at the lower end, but a long tail of rare, high values stretches to the right.
Intermittent Demand
A sales pattern where a product has zero demand for many periods, punctuated by occasional spikes.

Common questions

Why does Mean Absolute Error target the median?

MAE applies a linear penalty to errors, meaning it only cares about the absolute distance from the truth. Mathematically, the single value that minimizes the sum of absolute distances to all points in a dataset is the median, because it splits the data exactly in half.

When is it appropriate to use MAE?

MAE is ideal when the data is perfectly symmetrical, or when you are predicting a single continuous variable where extreme outliers should be ignored rather than chased. It is also useful when the cost of an error scales linearly with its size.

What is Tweedie deviance?

Tweedie deviance is a family of loss functions designed for right-skewed, positive-only data, such as insurance claims or retail sales. It allows forecasters to tune the penalty surface to account for the long tail without defaulting to a pure mean or median target.

Competing readings

Aggregate Forecasters

Prioritize the mean to ensure that item-level forecasts sum correctly to total volume.

For supply chain and inventory teams, aggregate accuracy is paramount. If a model systematically underforecasts because it targets the median, warehouses will not order enough stock to cover the long tail of demand spikes. These teams generally prefer MSE or asymmetric loss functions that penalize under-forecasting, ensuring that the sum of item-level predictions matches the true total demand.

Robustness Advocates

Prioritize the median to prevent extreme outliers from distorting the model's baseline predictions.

Engineers often favor MAE or Huber loss because these metrics are robust to extreme outliers. In datasets with corrupted records or massive anomalies, an MSE loss function will aggressively distort the model's weights to chase a single bad data point. They argue that it is better to train a stable model on MAE and accept the median target rather than letting outliers ruin the underlying regression.

Bias Correction Proponents

Advocate for training on MAE but mathematically adjusting the outputs to recover the mean.

This camp acknowledges the mathematical reality of MAE but refuses to abandon its stability. Instead of switching to MSE, they train models on MAE and apply post-hoc bias corrections, such as sales-weighting the training data. This forces the model to care more about high-volume days, dragging the median prediction closer to the true mean without exposing the core algorithm to quadratic outlier penalties.

Aggregate Forecasters 40%Robustness Advocates 35%Bias Correction Proponents 25%
Aggregate Forecasters
Prioritize the mean to ensure that item-level forecasts sum correctly to total volume.
Robustness Advocates
Prioritize the median to prevent extreme outliers from distorting the model's baseline predictions.
Bias Correction Proponents
Advocate for training on MAE but mathematically adjusting the outputs to recover the mean.

Perspectives this story doesn't cover

  • Business Stakeholders
  • Financial Planners

Sources

Source coverage

6 outlets

3 viewpoints surfaced

Aggregate Forecasters 40%Robustness Advocates 35%Bias Correction Proponents 25%
  1. [1]Albert Heijn TechnologyAggregate Forecasters

    The choice of loss function matters: MAE targets the median of the demand distribution

    Read on Albert Heijn Technology →
  2. [2]OpenForecastBias Correction Proponents

    Naughty APEs and the quest for the holy grail

    Read on OpenForecast →
  3. [3]Blue YonderRobustness Advocates

    How to evaluate Mean Absolute Error for probabilistic forecasts

    Read on Blue Yonder →
  4. [4]arXivBias Correction Proponents

    We also note that optimizing on MAE/ MAPE outputs median sales distribution

    Read on arXiv →
  5. [5]MetricGateRobustness Advocates

    MetricGate computes MAE and its companion diagnostics

    Read on MetricGate →
  6. [6]Factlen Editorial Team

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team →

Comments

Stay informed

Every angle. Every day.

Get Data & Analysis stories with full source coverage and perspective breakdowns, free every day.