Variable Selection Methods for Model-based Clustering
arXiv:1707.00306 · doi:10.1214/18-SS119
Abstract
Model-based clustering is a popular approach for clustering multivariate data which has seen applications in numerous fields. Nowadays, high-dimensional data are more and more common and the model-based clustering approach has adapted to deal with the increasing dimensionality. In particular, the development of variable selection techniques has received a lot of attention and research effort in recent years. Even for small size problems, variable selection has been advocated to facilitate the interpretation of the clustering results. This review provides a summary of the methods developed for variable selection in model-based clustering. Existing R packages implementing the different methods are indicated and illustrated in application to two data analysis examples.
References in corpus (10)
- Model-based clustering based on sparse finite Gaussian mixtures
- Simultaneous model-based clustering and visualization in the Fisher discriminative subspace
- Variable selection for model-based clustering using the integrated complete-data likelihood
- Variable Selection for Latent Class Analysis with Application to Low Back Pain Diagnosis
- Penalized model-based clustering with cluster-specific diagonal covariance matrices and grouped variables
- Model-based clustering using copulas with applications
- Model-based clustering for conditionally correlated categorical data
- Bayesian variable selection for latent class analysis using a collapsed Gibbs sampler
- clustvarsel: A Package Implementing Variable Selection for Model-based Clustering in R
- Variable selection for mixed data clustering: a model-based approach
Cited by in corpus (8)
- Simultaneous Dimension Reduction and Clustering via the NMF-EM Algorithm
- Simultaneous Dimensionality and Complexity Model Selection for Spectral Graph Clustering
- Clustering and Variable Selection in the Presence of Mixed Variable Types and Missing Data
- A sparse negative binomial mixture model for clustering RNA-seq count data
- High-dimensional clustering via Random Projections
- Model based clustering of multinomial count data
- A Bayesian Finite Mixture Model with Variable Selection for Data with Mixed-type Variables
- Model-based Clustering using Automatic Differentiation: Confronting Misspecification and High-Dimensional Data