Profile Likelihood Biclustering
arXiv:1206.6927 · doi:10.1214/19-EJS1667
Abstract
Biclustering, the process of simultaneously clustering the rows and columns of a data matrix, is a popular and effective tool for finding structure in a high-dimensional dataset. Many biclustering procedures appear to work well in practice, but most do not have associated consistency guarantees. To address this shortcoming, we propose a new biclustering procedure based on profile likelihood. The procedure applies to a broad range of data modalities, including binary, count, and continuous observations. We prove that the procedure recovers the true row and column classes when the dimensions of the data matrix tend to infinity, even if the functional form of the data distribution is misspecified. The procedure requires computing a combinatorial search, which can be expensive in practice. Rather than performing this search directly, we propose a new heuristic optimization procedure based on the Kernighan-Lin heuristic, which has nice computational properties and performs well in simulations. We demonstrate our procedure with applications to congressional voting records, and microarray analysis.
40 pages, 11 figures; R package in development at https://github.com/patperry/biclustpl
References in corpus (12)
- Modularity and community structure in networks
- Consistency of spectral clustering in stochastic block models
- Fast community detection by SCORE
- Pseudo-likelihood methods for community detection in large sparse networks
- Asymptotic normality of maximum likelihood and its variational approximation for stochastic blockmodels
- Convex Biclustering
- Optimal Estimation and Completion of Matrices with Biclustering Structures
- A Survey on Theoretical Advances of Community Detection in Networks
- Co-clustering separately exchangeable network data
- Belief propagation, robust reconstruction and optimal recovery of block models
- Convergence of the groups posterior distribution in latent or stochastic block models
- Matched bipartite block model with covariates