Bayesian Consensus Clustering
arXiv:1302.7280 · doi:10.1093/bioinformatics/btt425
Abstract
The task of clustering a set of objects based on multiple sources of data arises in several modern applications. We propose an integrative statistical model that permits a separate clustering of the objects for each data source. These separate clusterings adhere loosely to an overall consensus clustering, and hence they are not independent. We describe a computationally scalable Bayesian framework for simultaneous estimation of both the consensus clustering and the source-specific clusterings. We demonstrate that this flexible approach is more robust than joint clustering of all data sources, and is more powerful than clustering each data source separately. This work is motivated by the integrated analysis of heterogeneous biomedical data, and we present an application to subtype identification of breast cancer tumor samples using publicly available data from The Cancer Genome Atlas. Software is available at http://people.duke.edu/~el113/software.html.
32 pages, 13 figures
References in corpus (5)
- Proceedings of the 29th International Conference on Machine Learning (ICML-12)
- A simple example of Dirichlet process mixture inconsistency for the number of components
- Integrative Model-based clustering of microarray methylation and expression data
- Copula Mixture Model for Dependency-seeking Clustering
- Identifying cancer subtypes in glioblastoma by combining genomic, transcriptomic and epigenomic data
Cited by in corpus (25)
- A Survey on Multi-View Clustering
- Multiple kernel learning for integrative consensus clustering of 'omic datasets
- Integrative Generalized Convex Clustering Optimization and Feature Selection for Mixed Multi-View Data
- Escaping the curse of dimensionality in Bayesian model based clustering
- Shared kernel Bayesian screening
- Latent Simplex Position Model: High Dimensional Multi-view Clustering with Uncertainty Quantification
- D-GCCA: Decomposition-based Generalized Canonical Correlation Analysis for Multi-view High-dimensional Data
- Bayesian time-aligned factor analysis of paired multivariate time series
- Joint association and classification analysis of multi-view data
- Multi-Source Multi-View Clustering via Discrepancy Penalty
- Entropy Regularized Power k-Means Clustering
- CDPA: Common and Distinctive Pattern Analysis between High-dimensional Datasets
- Mutual Community Detection across Multiple Partially Aligned Social Networks
- A Dirichlet Process Mixture Model for Clustering Longitudinal Gene Expression Data
- Learning Sparsity and Block Diagonal Structure in Multi-View Mixture Models
- Scalable Heterogeneous Social Network Alignment through Synergistic Graph Partition
- LATTE: Application Oriented Social Network Embedding
- Bayesian Nonparametric Graph Clustering
- Directionally Dependent Multi-View Clustering Using Copula Model
- RaJIVE: Robust Angle Based JIVE for Integrating Noisy Multi-Source Data
- Angle-Based Joint and Individual Variation Explained
- Inferring Brain Signals Synchronicity from a Sample of EEG Readings
- Adaptive Weighted Multi-View Clustering
- Integrative clustering of high-dimensional data with joint and individual clusters, with an application to the Metabric study
- A Nonparametric Bayesian Method for Clustering of High-Dimensional Mixed Dataset