Rethinking Mean Square Error: Information, Generalized Estimation, and the James-Stein Paradox
arXiv:2412.08475
The paper compares mean square error with a newer ℓ‑information criterion for evaluating estimators, showing that while no pointwise risk like MSE has a uniformly optimal estimator, ℓ‑information is maximized by the score and explains why the James‑Stein estimator’s lower MSE does not contradict the optimality of maximum likelihood.
Abstract
The James-Stein estimator's dominance over maximum likelihood in mean square error has been called a paradox because maximum likelihood is known to be superior in many other respects. One response, due to Efron, is to question maximum likelihood. Another is to question MSE. We pursue the second and compare MSE with -information (Vos and Wu, 2025) as criteria for assessing estimators. The comparison rests on two distinctions: between point estimators and generalized estimators -- functions of the sample and parameter jointly, with the score as archetype -- as inferential objects, and between pointwise and family-aware assessment criteria. An elementary lemma shows that no pointwise criterion, MSE or any other risk built from a loss function, admits a uniformly optimal estimator; -information, which is family-aware and parameter-invariant, is uniformly maximized by the score. A point estimator is assessed through the generalized estimators it induces, and under the score map its -efficiency is the fraction of Fisher information the statistic retains, placing the criterion in Fisher's information-loss tradition. On unbiased estimators, -efficiency coincides with variance-based efficiency. Returning to James-Stein, the paradox dissolves: maximum likelihood is fully efficient because it is sufficient, while the James-Stein statistic is exactly two-to-one in the sample, and the information it destroys -- computed exactly -- is concentrated precisely where its MSE advantage is greatest. MSE retains its proper domain under genuine squared-error loss.
16 pages