Identifying Mixtures of Mixtures Using Bayesian Estimation
arXiv:1502.06449 · doi:10.1080/10618600.2016.1200472
Abstract
The use of a finite mixture of normal distributions in model-based clustering allows to capture non-Gaussian data clusters. However, identifying the clusters from the normal components is challenging and in general either achieved by imposing constraints on the model or by using post-processing procedures. Within the Bayesian framework we propose a different approach based on sparse finite mixtures to achieve identifiability. We specify a hierarchical prior where the hyperparameters are carefully selected such that they are reflective of the cluster structure aimed at. In addition this prior allows to estimate the model using standard MCMC sampling methods. In combination with a post-processing approach which resolves the label switching issue and results in an identified model, our approach allows to simultaneously (1) determine the number of clusters, (2) flexibly approximate the cluster distributions in a semi-parametric way using finite mixtures of normals and (3) identify cluster-specific parameters and classify observations. The proposed approach is illustrated in two simulation studies and on benchmark data sets.
49 pages
References in corpus (5)
- Natural Scales in Geographical Patterns
- Model-based clustering based on sparse finite Gaussian mixtures
- Parsimonious Skew Mixture Models for Model-Based Clustering and Classification
- Parsimonious Shifted Asymmetric Laplace Mixtures
- EMMIX-uskew: An R Package for Fitting Mixtures of Multivariate Skew t-distributions via the EM Algorithm
Cited by in corpus (12)
- Generalized mixtures of finite mixtures and telescoping sampling
- On the identifiability of Bayesian factor analytic models
- Bayesian Distance Clustering
- DS-UI: Dual-Supervised Mixture of Gaussian Mixture Models for Uncertainty Inference
- Finite mixture models do not reliably learn the number of components
- Clustering Multivariate Data using Factor Analytic Bayesian Mixtures with an Unknown Number of Components
- Better than the best? Answers via model ensemble in density-based clustering
- Estimating densities with nonlinear support using Fisher-Gaussian kernels
- Variational Inference and Sparsity in High-Dimensional Deep Gaussian Mixture Models
- Distributed Bayesian clustering using finite mixture of mixtures
- Deep mixture of linear mixed models for complex longitudinal data
- Coarsened mixtures of hierarchical skew normal kernels for flow cytometry analyses