Robust Mean Estimation on Highly Incomplete Data with Arbitrary Outliers
arXiv:2008.08071
Abstract
We study the problem of robustly estimating the mean of a -dimensional distribution given examples, where most coordinates of every example may be missing and examples may be arbitrarily corrupted. Assuming each coordinate appears in a constant factor more than examples, we show algorithms that estimate the mean of the distribution with information-theoretically optimal dimension-independent error guarantees in nearly-linear time . Our results extend recent work on computationally-efficient robust estimation to a more widely applicable incomplete-data setting.
29 pages, 2 figures. Published in AISTATS 2021. More details in the proof of Claim 14
References in corpus (11)
- Tuned Models of Peer Assessment in MOOCs
- Recent Advances in Algorithmic High-Dimensional Robust Statistics
- Robust Regression via Hard Thresholding
- Robust Hypothesis Testing Using Wasserstein Uncertainty Sets
- Principal Component Analysis with Contaminated Data: The High Dimensional Case
- High Dimensional Robust Sparse Regression
- Generalized Resilience and Robust Statistics
- Robust estimation via generalized quasi-gradients
- A Fast Spectral Algorithm for Mean Estimation with Sub-Gaussian Rates
- Learning Structured Distributions From Untrusted Batches: Faster and Simpler
- High-Dimensional Robust Mean Estimation via Gradient Descent