Mixtures of Generalized Hyperbolic Distributions and Mixtures of Skew-t Distributions for Model-Based Clustering with Incomplete Data
arXiv:1703.02177 · doi:10.1016/j.csda.2018.08.016
Abstract
Robust clustering from incomplete data is an important topic because, in many practical situations, real data sets are heavy-tailed, asymmetric, and/or have arbitrary patterns of missing observations. Flexible methods and algorithms for model-based clustering are presented via mixture of the generalized hyperbolic distributions and its limiting case, the mixture of multivariate skew-t distributions. An analytically feasible EM algorithm is formulated for parameter estimation and imputation of missing values for mixture models employing missing at random mechanisms. The proposed methodologies are investigated through a simulation study with varying proportions of synthetic missing values and illustrated using a real dataset. Comparisons are made with those obtained from the traditional mixture of generalized hyperbolic distribution counterparts by filling in the missing data using the mean imputation method.
References in corpus (4)
Cited by in corpus (6)
- A Comparative Study of Methods for Estimating Conditional Shapley Values and When to Use Them
- A Robust and Flexible EM Algorithm for Mixtures of Elliptical Distributions with Missing Data
- Fast model-based clustering of partial records
- A Bayesian approach for clustering skewed data using mixtures of multivariate normal-inverse Gaussian distributions
- Infinite mixtures of multivariate normal-inverse Gaussian distributions for clustering of skewed data
- Flexible High-Dimensional Unsupervised Learning with Missing Data