High-dimensional regression with noisy and missing data: Provable guarantees with nonconvexity
arXiv:1109.3714 · doi:10.1214/12-AOS1018
Abstract
Although the standard formulations of prediction problems involve fully-observed and noiseless data drawn in an i.i.d. manner, many applications involve noisy and/or missing data, possibly involving dependence, as well. We study these issues in the context of high-dimensional sparse linear regression, and propose novel estimators for the cases of noisy, missing and/or dependent data. Many standard approaches to noisy or missing data, such as those using the EM algorithm, lead to optimization problems that are inherently nonconvex, and it is difficult to establish theoretical guarantees on practical algorithms. While our approach also involves optimizing nonconvex programs, we are able to both analyze the statistical error associated with any global optimum, and more surprisingly, to prove that a simple algorithm based on projected gradient descent will converge in polynomial time to a small neighborhood of the set of all global minimizers. On the statistical side, we provide nonasymptotic bounds that hold with high probability for the cases of noisy, missing and/or dependent data. On the computational side, we prove that under the same types of conditions required for statistical consistency, the projected gradient descent algorithm is guaranteed to converge at a geometric rate to a near-global minimizer. We illustrate these theoretical predictions with simulations, showing close agreement with the predicted scalings.
Published in at http://dx.doi.org/10.1214/12-AOS1018 the Annals of Statistics (http://www.imstat.org/aos/) by the Institute of Mathematical Statistics (http://www.imstat.org)
References in corpus (3)
Cited by in corpus (162)
- Challenges of Big Data Analysis
- Regularized estimation in sparse high-dimensional time series models
- Compressed Sensing using Generative Models
- Optimal detection of sparse principal components in high dimension
- Complete Dictionary Recovery over the Sphere I: Overview and the Geometric Picture
- Robust subspace clustering
- Signal Processing on Graphs: Causal Modeling of Unstructured Data
- Regularized M-estimators with nonconvexity: Statistical and algorithmic theory for local optima
- Compressed Sensing with Deep Image Prior and Learned Regularization
- Low-complexity Multiclass Encryption by Compressed Sensing
- -penalized maximum likelihood for sparse directed acyclic graphs
- High-dimensional learning of linear causal networks via inverse covariance estimation
- A Direct Estimation of High Dimensional Stationary Vector Autoregressions
- Structure estimation for discrete graphical models: Generalized covariance matrices and their inverses
- On Known-Plaintext Attacks to a Compressed Sensing-based Encryption: A Quantitative Analysis
- Confidence sets in sparse regression
- Semi-Blind Inference of Topologies and Dynamical Processes over Graphs
- On Iterative Hard Thresholding Methods for High-dimensional M-Estimation
- Moving Beyond Sub-Gaussianity in High-Dimensional Statistics: Applications in Covariance Estimation and Linear Regression
- Local Privacy, Data Processing Inequalities, and Statistical Minimax Rates
- Learning partial differential equations for biological transport models from noisy spatiotemporal data
- The Landscape of Empirical Risk for Non-convex Losses
- A Tight Bound of Hard Thresholding
- Measurement Error in Lasso: Impact and Correction
- Regularized EM Algorithms: A Unified Framework and Statistical Guarantees
- Noisy Sparse Subspace Clustering
- Sparse Signal Processing with Linear and Nonlinear Observations: A Unified Shannon-Theoretic Approach
- Zeroth Order Nonconvex Multi-Agent Optimization over Networks
- NESTT: A Nonconvex Primal-Dual Splitting Method for Distributed and Stochastic Optimization
- Distributed Robust Learning
- Iterative Hessian sketch: Fast and accurate solution approximation for constrained least-squares
- On the Adversarial Robustness of LASSO Based Feature Selection
- Statistical Inference for Model Parameters in Stochastic Gradient Descent
- A unifying approach for doubly-robust regularized estimation of causal contrasts
- Covariate Selection in High-Dimensional Generalized Linear Models With Measurement Error
- Efficient Algorithms for Outlier-Robust Regression
- FNETS: Factor-adjusted network estimation and forecasting for high-dimensional time series
- Estimating Structured Vector Autoregressive Model
- Linear Regression with Limited Observation
- Fast and Robust Least Squares Estimation in Corrupted Linear Models
- On Robustness of Principal Component Regression
- Localizing Changes in High-Dimensional Vector Autoregressive Processes
- Regularized Estimation and Testing for High-Dimensional Multi-Block Vector-Autoregressive Models
- Learning with Non-Convex Truncated Losses by SGD
- Penalized Maximum Likelihood Estimation of Multi-layered Gaussian Graphical Models
- Penalized Estimation and Forecasting of Multiple Subject Intensive Longitudinal Data
- Covariance Matrix Estimation with Non Uniform and Data Dependent Missing Observations
- The estimation error of general first order methods
- De-biased sparse PCA: Inference and testing for eigenstructure of large covariance matrices
- Sampling Requirements for Stable Autoregressive Estimation
- Pivotal Estimation via Self-Normalization for High-Dimensional Linear Models with Error in Variables
- Robust High Dimensional Sparse Regression and Matching Pursuit
- High-Dimensional Multivariate Time Series With Additional Structure
- Greedy algorithms for prediction
- Orthogonal Matching Pursuit with Noisy and Missing Data: Low and High Dimensional Results
- High-dimensional Linear Discriminant Analysis: Optimality, Adaptive Algorithm, and Missing Data
- Rejoinder: Latent variable graphical model selection via convex optimization
- Optimal Algorithms for Ridge and Lasso Regression with Partially Observed Attributes
- Non-Asymptotic Guarantees for Reliable Identification of Granger Causality via the LASSO
- A Likelihood Ratio Framework for High Dimensional Semiparametric Regression
- Two Stage Non-penalized Corrected Least Squares for High Dimensional Linear Models with Measurement error or Missing Covariates
- High Dimensional M-Estimation with Missing Outcomes: A Semi-Parametric Framework
- Preserving Differential Privacy Between Features in Distributed Estimation
- High dimensional errors-in-variables models with dependent measurements
- Joint Estimation and Inference for Data Integration Problems based on Multiple Multi-layered Gaussian Graphical Models
- Valid Post-selection Inference in Assumption-lean Linear Regression
- Inference for Heteroskedastic PCA with Missing Data
- Differentially Private (Gradient) Expectation Maximization Algorithm with Statistical Guarantees
- Regularization Approach for Network Modeling of German Power Derivative Market
- Random Forest Missing Data Algorithms
- Subspace Estimation from Unbalanced and Incomplete Data Matrices: Statistical Guarantees
- Graph quilting: graphical model selection from partially observed covariances
- Online and Distributed Robust Regressions under Adversarial Data Corruption
- Penalized pairwise pseudo likelihood for variable selection with nonignorable missing data
- Compressive Coded Aperture Keyed Exposure Imaging with Optical Flow Reconstruction
- Minimum Distance Estimation for Robust High-Dimensional Regression
- Functional Linear Regression with Mixed Predictors
- CoCoLasso for High-dimensional Error-in-variables Regression
- Hybrid Modeling of Regional COVID-19 Transmission Dynamics in the U.S
- Sparse Recovery with Linear and Nonlinear Observations: Dependent and Noisy Data
- Learning Ising Models with Independent Failures
- Finite Sample Theory for High-Dimensional Functional/Scalar Time Series with Applications
- Linear convergence of SDCA in statistical estimation
- Multi-linear Tensor Autoregressive Models
- Sparse Linear Regression With Missing Data
- The correlation-assisted missing data estimator
- A Survey of Estimation Methods for Sparse High-dimensional Time Series Models
- Missing Data in Sparse Transition Matrix Estimation for Sub-Gaussian Vector Autoregressive Processes
- Nonparametric covariance estimation for mixed longitudinal studies, with applications in midlife women's health
- An Investigation of Methods for Handling Missing Data with Penalized Regression
- An -Regularization Approach to High-Dimensional Errors-in-variables Models
- Statistical Analysis of Stationary Solutions of Coupled Nonconvex Nonsmooth Empirical Risk Minimization
- SAGA and Restricted Strong Convexity
- Non-separable covariance models for spatio-temporal data, with applications to neural encoding analysis
- High-Dimensional Robust Mean Estimation via Gradient Descent
- Regularized deep learning with nonconvex penalties
- Multi-Task Learning with Incomplete Data for Healthcare
- Sparse Principal Component Analysis for High Dimensional Vector Autoregressive Models
- Robust High Dimensional Expectation Maximization Algorithm via Trimmed Hard Thresholding
- Non-bifurcating phylogenetic tree inference via the adaptive LASSO
- Joint Estimation of Multiple Graphical Models from High Dimensional Time Series
- Scalable Interpretable Learning for Multi-Response Error-in-Variables Regression
- Implicit Regularization and Entrywise Convergence of Riemannian Optimization for Low Tucker-Rank Tensor Completion
- Estimating Network Structure from Incomplete Event Data
- High-dimensional Log-Error-in-Variable Regression with Applications to Microbial Compositional Data Analysis
- Methods for Recovering Conditional Independence Graphs: A Survey
- On Learning Sparsely Used Dictionaries from Incomplete Samples
- Understanding Notions of Stationarity in Non-Smooth Optimization
- Detection and estimation of parameters in high dimensional multiple change point regression models via regularization and discrete optimization
- Minimax Rate-optimal Estimation of High-dimensional Covariance Matrices with Incomplete Data
- Confidence intervals for parameters in high-dimensional sparse vector autoregression
- A Fast Detection Method of Break Points in Effective Connectivity Networks
- Inference on the Change Point for High Dimensional Dynamic Graphical Models
- Robust Variable Selection under Cellwise Contamination
- Convex and Non-convex Approaches for Statistical Inference with Class-Conditional Noisy Labels
- Rate Optimal Estimation and Confidence Intervals for High-dimensional Regression with Missing Covariates
- Inference in High-Dimensional Linear Measurement Error Models
- Least Squares with Error in Variables
- High-Dimensional Covariance Decomposition into Sparse Markov and Independence Domains
- High-Dimensional Covariance Decomposition into Sparse Markov and Independence Models
- Debiasing Stochastic Gradient Descent to handle missing values
- Causal Inference with Corrupted Data: Measurement Error, Missing Values, Discretization, and Differential Privacy
- Low-rank matrix estimation in multi-response regression with measurement errors: Statistical and computational guarantees
- New Computational and Statistical Aspects of Regularized Regression with Application to Rare Feature Selection and Aggregation
- Outlier-robust sparse/low-rank least-squares regression and robust matrix completion
- Learning the Dynamics of Sparsely Observed Interacting Systems
- Selective Inference and Learning Mixed Graphical Models
- High-Dimensional Semiparametric Selection Models: Estimation Theory with an Application to the Retail Gasoline Market
- Tuning-Free Heterogeneity Pursuit in Massive Networks
- Analysis of High Dimensional Compositional Data Containing Structural Zeros with Applications to Microbiome Data
- Direct estimation and inference of differential Granger causality between two high-dimensional time series
- Sparse spectral estimation with missing and corrupted measurements
- Prediction in the presence of response-dependent missing labels
- Approximate Factor Models with Strongly Correlated Idiosyncratic Errors
- A Simple Correction Procedure for High-Dimensional Generalized Linear Models with Measurement Error
- Calibrated zero-norm regularized LS estimator for high-dimensional error-in-variables regression
- Adaptive Estimation and Statistical Inference for High-Dimensional Graph-Based Linear Models
- A General Theory for Large-Scale Curve Time Series via Functional Stability Measure
- Tensor models for linguistics pitch curve data of native speakers of Afrikaans
- Sharp global convergence guarantees for iterative nonconvex optimization: A Gaussian process perspective
- High-dimensional Simultaneous Inference on Non-Gaussian VAR Model via De-biased Estimator
- Robust Inference of Risks of Large Portfolios
- Uniform-in-Submodel Bounds for Linear Regression in a Model Free Framework
- Contextual Linear Bandits under Noisy Features: Towards Bayesian Oracles
- Inference on the change point with the jump size near the boundary of the region of detectability in high dimensional time series models
- Errors-in-variables models with dependent measurements
- Minimax rates of -losses for high-dimensional linear regression models with additive measurement errors over -balls
- Kernel Ordinary Differential Equations
- Robust Elastic Net Regression
- Generalization Bounds for High-dimensional M-estimation under Sparsity Constraint
- A Bernstein-type Inequality for High Dimensional Linear Processes with Applications to Robust Estimation of Time Series Regressions
- A Robust Time Series Model with Outliers and Missing Entries
- Robust Compressed Sensing Under Matrix Uncertainties
- Robust Regression via Online Feature Selection under Adversarial Data Corruption
- On the uniform convergence of empirical norms and inner products, with application to causal inference
- Estimation Rates for Sparse Linear Cyclic Causal Models
- Multi-task Learning with High-Dimensional Noisy Images
- Robust Lasso with missing and grossly corrupted observations
- Model-Assisted Uniformly Honest Inference for Optimal Treatment Regimes in High Dimension
- Minimax Estimation of Partially-Observed Vector AutoRegressions
- HMLasso: Lasso with High Missing Rate
- Logistic regression and Ising networks: prediction and estimation when violating lasso assumptions