Weak Moment Methods for Statistical Inference: with an Application to Robust Estimation
arXiv:2604.23619
Abstract
A companion paper develops a generalised framework in which a probability law is represented by a tempered distribution - on the same footing as a density or characteristic function - and information is extracted by pairing with a positive Schwartz kernel that acts as a measurement instrument rather than as part of the law; the resulting weak moments of all orders exist unconditionally. The present paper turns this into a methodology for statistical inference: estimation via weak moment matching, weak characteristic functions, weak cumulants, and regularised density reconstruction by Tikhonov inversion. Parametric inference proceeds directly from weak expectations, without reconstructing the density. The central result is that weak moment estimators are automatically locally robust in the sense of Hampel: their score is bounded and redescending, their influence function has a closed form, and their gross error sensitivity is finite in every identifiable parametric model - all inherited from the kernel's decay, with no ad hoc truncation. The kernel plays the role of Huber's tuning constant, but as a structural component of the model rather than a post-hoc modification. The framework is worked out for the Cauchy location model (where no classical moment estimator exists), a Student location-scale model, a bivariate Cauchy location model, a bivariate location-scale model, and a location of a moving atom (a non-dominated model). Monte Carlo comparisons show weak moment estimators matching or outperforming classical robust benchmarks under contamination; in the bivariate case the MLE scale estimate breaks down while the weak moment estimator converges at the parametric rate. The reconstruction route is inherently non-parametric and opens a path to weak density estimation.
34 pages, one figure and five tables. Inserted further detail and additional examples. The distributional representation now is given by a tempered distribution instead of a pair formed by a tempered distribution and a kernel) and the kernel is viewed as an instrument for extracting information off the probability law from the tempered distribution