Operator norm consistent estimation of large-dimensional sparse covariance matrices
arXiv:0901.3220 · doi:10.1214/07-AOS559
Abstract
Estimating covariance matrices is a problem of fundamental importance in multivariate statistics. In practice it is increasingly frequent to work with data matrices of dimension , where and are both large. Results from random matrix theory show very clearly that in this setting, standard estimators like the sample covariance matrix perform in general very poorly. In this "large , large " setting, it is sometimes the case that practitioners are willing to assume that many elements of the population covariance matrix are equal to 0, and hence this matrix is sparse. We develop an estimator to handle this situation. The estimator is shown to be consistent in operator norm, when, for instance, we have as . In other words the largest singular value of the difference between the estimator and the population covariance matrix goes to zero. This implies consistency of all the eigenvalues and consistency of eigenspaces associated to isolated eigenvalues. We also propose a notion of sparsity for matrices, that is, "compatible" with spectral analysis and is independent of the ordering of the variables.
Published in at http://dx.doi.org/10.1214/07-AOS559 the Annals of Statistics (http://www.imstat.org/aos/) by the Institute of Mathematical Statistics (http://www.imstat.org)
References in corpus (2)
Cited by in corpus (14)
- Covariance regularization by thresholding
- Sparse permutation invariant covariance estimation
- Optimal rates of convergence for covariance matrix estimation
- Covariance Estimation: The GLM and Regularization Perspectives
- Optimal rates of convergence for sparse covariance matrix estimation
- Vast volatility matrix estimation for high-frequency financial data
- Adaptive covariance matrix estimation through block thresholding
- High-dimensional analysis of semidefinite relaxations for sparse principal components
- High-dimensionality effects in the Markowitz problem and other quadratic programs with linear constraints: Risk underestimation
- Regression on manifolds: Estimation of the exterior derivative
- On information plus noise kernel random matrices
- A penalized empirical likelihood method in high dimensions
- Convergence of the largest eigenvalue of normalized sample covariance matrices when p and n both tend to infinity with their ratio converging to zero
- Discussion of: Treelets--An adaptive multi-scale basis for sparse unordered data