Model Based Clustering of High-Dimensional Binary Data
arXiv:1404.3174 · doi:10.1016/j.csda.2014.12.009
Abstract
We propose a mixture of latent trait models with common slope parameters (MCLT) for model-based clustering of high-dimensional binary data, a data type for which few established methods exist. Recent work on clustering of binary data, based on a -dimensional Gaussian latent variable, is extended by incorporating common factor analyzers. Accordingly, our approach facilitates a low-dimensional visual representation of the clusters. We extend the model further by the incorporation of random block effects. The dependencies in each block are taken into account through block-specific parameters that are considered to be random variables. A variational approximation to the likelihood is exploited to derive a fast algorithm for determining the model parameters. Our approach is demonstrated on real and simulated data.
References in corpus (6)
- Natural Scales in Geographical Patterns
- Mixtures of Skew-t Factor Analyzers
- Parsimonious Skew Mixture Models for Model-Based Clustering and Classification
- Estimating Common Principal Components in High Dimensions
- Variational Bayes Approximations for Clustering via Mixtures of Normal Inverse Gaussian Distributions
- Mixtures of Common Skew-t Factor Analyzers
Cited by in corpus (5)
- Mixtures of Skew-t Factor Analyzers
- Conditionally conjugate mean-field variational Bayes for logistic models
- A Bayesian non-parametric method for clustering high-dimensional binary data
- A pairwise likelihood approach to simultaneous clustering and dimensional reduction of ordinal data
- Flexible High-Dimensional Unsupervised Learning with Missing Data