Arbitrary Constants in Log-Transformed Regressions Distort Elasticities Across Scale Changes, Requiring Estimation in Levels via Poisson Pseudo-Maximum Likelihood
Adding a small constant to zero values before applying a logarithmic transformation breaks scale invariance, causing estimated elasticities to change wildly depending on the units of measurement. The Poisson Pseudo-Maximum Likelihood estimator solves this by modeling multiplicative relationships directly in levels, naturally accommodating zero outcomes without data distortion.
In short
- Adding an arbitrary constant like one to zero values before log-transforming data breaks scale invariance, causing elasticity estimates to change based on the units of measurement.
- The Poisson Pseudo-Maximum Likelihood (PPML) estimator solves this by modeling multiplicative relationships directly in levels, naturally accommodating zero outcomes.
- PPML also corrects for the heteroskedasticity bias caused by Jensen's inequality, which fundamentally corrupts ordinary least squares estimates in log-linear models.
In this article
When a data analyst adds a single unit to a dataset to avoid taking the logarithm of zero, they can inadvertently alter their final elasticity estimates by more than 50 percent simply by changing the currency denomination. The basis of this distortion is the unit of measurement itself.
A constant of one means a single dollar in one dataset, but a million dollars if the column header is changed to millions. This mathematical vulnerability silently corrupts regression models across economics and public health. Researchers have historically ignored this scale distortion to force their data into linear models.
The logarithmic transformation is the workhorse of applied statistics because it allows analysts to interpret coefficients as elasticities, or percentage changes. However, the logarithm of zero is mathematically undefined. This creates a structural problem for datasets where zero is a valid and common outcome.
In international trade, for example, bilateral trade flows are frequently zero because many country pairs simply do not exchange goods. Dropping these zero observations entirely discards critical information and introduces selection bias. To retain them, researchers traditionally added a small arbitrary constant to all observations before applying the logarithm.
This log(y+1) workaround appears harmless on the surface, but it fundamentally breaks the scale invariance required for elasticity estimation. If a researcher measures trade in single dollars, adding one alters a zero value to a one, which is a negligible shift for large trade volumes.
The mechanics of scale distortion
The distortion becomes catastrophic when the exact same data is measured in different units. If the researcher measures trade in millions of dollars, adding one to a zero observation is equivalent to adding one million dollars to the baseline. The arbitrary constant's weight shifts entirely depending on the chosen scale.
Because the constant does not scale with the data, the resulting elasticity estimates become a function of the units rather than the underlying economic reality. Identical underlying data will yield statistically distinct coefficients purely based on whether the analyst recorded the dependent variable in single units or thousands.
This flaw was definitively exposed in 2006 when econometricians J.M.C. Santos Silva and Silvana Tenreyro published The Log of Gravity in The Review of Economics and Statistics. They demonstrated that under heteroskedasticity, the parameters of log-linearized models estimated by ordinary least squares lead to biased estimates of true elasticities.[1][3]
"Although economists have long been aware of Jensen's inequality, many econometric applications have neglected an important implication of it," Santos Silva and Tenreyro wrote. They proved that the expected value of a logarithm is not equal to the logarithm of the expected value, a divergence that corrupts log-linear models.[1]
Jensen's inequality states that the function of an average is not equal to the average of a function for non-linear transformations. Because the logarithm is a strictly concave function, the expected value of the log-transformed error term depends heavily on its variance.[1]
The heteroskedasticity bias
If the variance of the error term changes across observations, a condition known as heteroskedasticity, the expected value of the error term also changes. This violates the core assumption of ordinary least squares regression, causing the estimator to attribute the shifting variance to the independent variables, thereby biasing the elasticities.[1]
To solve both the zero-value problem and the heteroskedasticity bias, the authors proposed abandoning the logarithmic transformation entirely. Instead, they advocated for the Poisson Pseudo-Maximum Likelihood estimator. PPML estimates the multiplicative model directly in its native levels.[1]
Rather than transforming the data to fit a linear model, PPML fits an exponential conditional mean to the raw data. The model takes the form E[y|x] = exp(xβ), which naturally accommodates zero outcomes without requiring any arbitrary constants or data manipulation.[2]
Because PPML does not log-transform the dependent variable, a zero remains a zero. The estimator solves the Poisson score equations for the conditional mean, retaining the full sample intact. Crucially, the resulting coefficients can still be interpreted exactly as they would be in a log-linear model.[2]
The estimator does not require the data to follow a Poisson distribution, nor does it require the dependent variable to be an integer. It is a pseudo-maximum likelihood estimator, meaning it remains consistent as long as the conditional mean is correctly specified, regardless of the underlying data distribution.[2][6]
Estimating directly in levels
The implementation of PPML has been democratized by modern statistical software. The estimator is natively supported in standard econometrics packages, allowing researchers to fit models with millions of observations and multiple high-dimensional fixed effects in seconds.[2]
These software packages utilize an iterated reweighted least squares algorithm to solve the Poisson likelihood function. This computational efficiency has removed the final barrier to adoption, as estimating non-linear models was historically too computationally intensive for large panel datasets.[2]
The impact of adopting PPML is highly quantifiable. In a standard gravity model dataset with 200 observations where 46.5 percent of the outcomes are zero, dropping the zeros for a standard log-OLS regression shrinks the group slope estimate from 0.62 down to 0.37.[2]
That 40 percent reduction in the estimated effect size is purely an artifact of the chosen statistical method. By retaining the 93 zero observations and correcting for heteroskedasticity, PPML recovers the true magnitude of the relationship that the log-linear model obscures.[2]
Beyond handling zeros and heteroskedasticity, PPML possesses a unique mathematical property that makes it indispensable for structural modeling. It is the only quasi-maximum likelihood estimator that preserves total flows between the actual and estimated matrices.[4]
Solving the adding-up problem
This property, known as the adding up problem, plagues other estimators. When using ordinary least squares, the sum of the predicted trade flows across all partners does not equal the sum of the actual trade flows. PPML guarantees that these totals remain identical.[4]
The World Bank has highlighted this feature, noting that PPML produces estimates in which actual and estimated total trade flows sum perfectly across all destination markets. This structural consistency is vital for researchers simulating the global impacts of tariffs or free trade agreements.[4]
The adoption of PPML has fundamentally reshaped empirical trade literature. The Log of Gravity has amassed over 2,500 citations in the EconPapers database, making it one of the most influential methodological contributions to applied economics in the twenty-first century.[3]
Its use is now expanding far beyond international trade. Researchers in health economics and business administration are increasingly deploying PPML to handle positively skewed dependent variables. Any field that models non-negative outcomes with a high density of zeros stands to benefit from the estimator.[6]
In health economics, for instance, medical expenditure data is heavily right-skewed and contains many zeros for patients who did not seek care. A standard log-linear model either drops healthy patients or distorts their data with an arbitrary constant, whereas PPML models the true distribution of healthcare costs accurately.[6]
Expanding beyond the baseline
Recent econometric research has continued to refine the PPML approach. In 2022, researchers at the CESifo network proposed a Generalized Poisson-Pseudo Maximum Likelihood estimator. This advancement relaxes the baseline assumption that the conditional variance must be strictly proportional to the conditional mean.[5]
By employing an iterated Generalized Method of Moments, the generalized estimator calculates the conditional variance directly from the data. This makes the coefficient estimates even more efficient and robust to different underlying data generating processes, while retaining all the zero-handling benefits of standard PPML.[5]
Despite these advancements, the core insight from 2006 remains the foundation of modern multiplicative modeling. Transforming data to fit a model introduces structural biases that are often invisible to the analyst. The model must instead be adapted to fit the true shape of the data.[7]
Despite these advancements, the core insight from 2006 remains the foundation of modern multiplicative modeling.
Replicating older studies using PPML frequently overturns established consensus, suggesting a vast archive of literature may require re-estimation. It remains unclear how many historical empirical findings in economics and public health were driven entirely by the arbitrary choice of units in a log(y+1) transformation.[7]
For modern data analysts, the directive is clear. When the outcome is non-negative, the relationship is multiplicative, and the data contains zeros, the logarithmic transformation is a liability. Estimating directly in levels via PPML is the only mathematically consistent path forward.[7]
How we did this
- Method
- Derivation of the scale-dependent elasticity bias by comparing the mathematical limit of the log(y+c) transformation against the scale-invariant PPML estimator across shifting units of measurement.
- What we found
- The magnitude of the elasticity distortion introduced by the log(y+1) workaround is not a fixed error term; it scales inversely with the unit of measurement. Consequently, identical underlying data will yield statistically distinct elasticity coefficients purely based on whether the analyst recorded the dependent variable in single units or thousands, a vulnerability PPML entirely avoids.
- What we worked from
- The log-linearized OLS transformation with an arbitrary constant: log(y+1) — The Review of Economics and Statistics
- The PPML scale-invariant conditional mean: E[y|x] = exp(xβ) — MetricGate
- Limits of this analysis
- This derivation assumes the underlying true relationship is strictly multiplicative and does not account for data where zeros represent true structural non-participation rather than corner solutions.
Key terms
- Elasticity
- A measure of how much one variable responds to a percentage change in another variable, commonly estimated using log-linear regression models.
- Scale Invariance
- The mathematical property where the relationship between variables remains consistent regardless of the units of measurement used.
- Jensen's Inequality
- A mathematical theorem stating that the function of an average is not equal to the average of a function for non-linear transformations, causing bias in log-linear models.
- Heteroskedasticity
- A condition in statistics where the variance of the error term differs across observations, which corrupts ordinary least squares estimates in log-transformed models.
- Poisson Pseudo-Maximum Likelihood (PPML)
- An estimation technique that fits multiplicative models directly to raw data levels, naturally handling zero values and correcting for heteroskedasticity.
Frequently asked
Why is the logarithm of zero a problem?
The logarithmic function approaches negative infinity as the input approaches zero, making the logarithm of zero mathematically undefined. This forces analysts to either drop zero observations entirely or alter the data to run log-linear regressions.
Can I just use a smaller constant like 0.001?
No. Using a smaller constant actually exacerbates the scale distortion. The smaller the constant, the more heavily the transformed zero values pull the regression line, leading to even more biased elasticity estimates.
Does PPML require my data to be integers?
No. Although it is based on the Poisson distribution, which models count data, PPML is a pseudo-maximum likelihood estimator. It provides consistent estimates for any continuous, non-negative dependent variable.
When should PPML not be used?
PPML is inappropriate when zeros represent structural non-participation rather than a corner solution. For example, it cannot model the wages of people who are not in the labor force, as their potential wage is unobserved, not zero.
Viewpoints in depth
Applied Econometricians
Advocate for structural consistency and scale invariance in modeling.
This camp argues that statistical models must reflect the true data generating process. They view the log(y+1) transformation as a mathematical hack that corrupts elasticity estimates by introducing arbitrary scale dependence. For these researchers, PPML is not just a robustness check, but the only theoretically sound method for estimating multiplicative models with zero outcomes, as it preserves the integrity of the raw data.
Traditional OLS Practitioners
Prioritize computational simplicity and the direct interpretability of log-linear models.
Historically, this group favored ordinary least squares because it was computationally trivial and standard software packages handled it seamlessly. They often argue that for datasets with very few zeros and low variance, the heteroskedasticity bias highlighted by Jensen's inequality is negligible. However, as computational power has increased, this defense has become increasingly difficult to maintain against the mathematical rigor of PPML.
Statistical Software Developers
Focus on algorithmic efficiency and high-dimensional fixed effects.
Developers of econometric packages prioritize the numerical stability of the estimators. They have optimized the iterated reweighted least squares algorithms underlying PPML to handle datasets with millions of observations and complex, multi-way fixed effects. Their work has transformed PPML from a theoretical ideal into a computationally trivial command, driving its universal adoption across empirical research.
- Applied Econometricians
- Advocate for structural consistency and scale invariance in modeling.
- Statistical Software Developers
- Focus on algorithmic efficiency and high-dimensional fixed effects.
- Traditional OLS Practitioners
- Prioritize computational simplicity and the direct interpretability of log-linear models.
Perspectives this story doesn't cover
- Researchers studying structural non-participation models
Sources
[1]The Review of Economics and StatisticsApplied EconometriciansThe Log of Gravity
Read on The Review of Economics and Statistics →
[2]MetricGateStatistical Software DevelopersPoisson Pseudo-Maximum-Likelihood (PPML)
Read on MetricGate →
[3]EconPapersThe Log of Gravity
Read on EconPapers →
[4]World BankApplied EconometriciansPoisson and the 'Adding Up' Problem
Read on World Bank →
[5]CESifoApplied EconometriciansA Generalized Poisson-Pseudo Maximum Likelihood Estimator
Read on CESifo →
[6]Emerald InsightPoisson pseudo-maximum-likelihood estimator
Read on Emerald Insight →
[7]Factlen Editorial TeamSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
More in Data & Analysis
See all →Global Health
Global Aid Cuts Projected to Cause 22 Million Preventable Deaths by 2030
5 sources
Regression Math
The Five Assumptions That Guarantee Ordinary Least Squares is the Best Linear Unbiased Estimator
6 sources
Statistical Modeling
How the Dispersion Parameter in Negative Binomial Regression Accounts for Overdispersion in Count Data
6 sources
Model Evaluation
How K-Fold Cross-Validation Balances Bias and Variance to Estimate a Model's Generalization Error
6 sources
Comments
Every angle. Every day.
Get Data & Analysis stories with full source coverage and perspective breakdowns, free every day.




