Statistical Significance of Clustering using Soft Thresholding
arXiv:1305.5879 · doi:10.1080/10618600.2014.948179
Abstract
Clustering methods have led to a number of important discoveries in bioinformatics and beyond. A major challenge in their use is determining which clusters represent important underlying structure, as opposed to spurious sampling artifacts. This challenge is especially serious, and very few methods are available, when the data are very high in dimension. Statistical Significance of Clustering (SigClust) is a recently developed cluster evaluation tool for high dimensional low sample size data. An important component of the SigClust approach is the very definition of a single cluster as a subset of data sampled from a multivariate Gaussian distribution. The implementation of SigClust requires the estimation of the eigenvalues of the covariance matrix for the null multivariate Gaussian distribution. We show that the original eigenvalue estimation can lead to a test that suffers from severe inflation of type-I error, in the important case where there are a few very large eigenvalues. This paper addresses this critical challenge using a novel likelihood based soft thresholding approach to estimate these eigenvalues, which leads to a much improved SigClust. Major improvements in SigClust performance are shown by both mathematical analysis, based on the new notion of Theoretical Cluster Index, and extensive simulation studies. Applications to some cancer genomic data further demonstrate the usefulness of these improvements.
27 pages, 5 figures
References in corpus (5)
- High-dimensional graphs and variable selection with the Lasso
- Sparse permutation invariant covariance estimation
- Network exploration via the adaptive LASSO and SCAD penalties
- Penalized model-based clustering with cluster-specific diagonal covariance matrices and grouped variables
- A Constrained L1 Minimization Approach to Sparse Precision Matrix Estimation
Cited by in corpus (5)
- On perfect clustering of high dimension, low sample size data
- Rank-transformed subsampling: inference for multiple data splitting and exchangeable p-values
- Gaussian Mixture Clustering Using Relative Tests of Fit
- New asteroid clusters and evidence of collisional fragmentation in the L5 Trojan cloud of Mars
- Large dimensional analysis of general margin based classification methods