Clustering, Classification, Discriminant Analysis, and Dimension Reduction via Generalized Hyperbolic Mixtures
arXiv:1308.6315 · doi:10.1016/j.csda.2015.10.008
Abstract
A method for dimension reduction with clustering, classification, or discriminant analysis is introduced. This mixture model-based approach is based on fitting generalized hyperbolic mixtures on a reduced subspace within the paradigm of model-based clustering, classification, or discriminant analysis. A reduced subspace of the data is derived by considering the extent to which group means and group covariances vary. The members of the subspace arise through linear combinations of the original data, and are ordered by importance via the associated eigenvalues. The observations can be projected onto the subspace, resulting in a set of variables that captures most of the clustering information available. The use of generalized hyperbolic mixtures gives a robust framework capable of dealing with skewed clusters. Although dimension reduction is increasingly in demand across many application areas, the authors are most familiar with biological applications and so two of the five real data examples are within that sphere. Simulated data are also used for illustration. The approach introduced herein can be considered the most general such approach available, and so we compare results to three special and limiting cases. Comparisons with several well established techniques illustrate its promising performance.
References in corpus (13)
- In silico prediction of protein-protein interactions in human macrophages
- Natural Scales in Geographical Patterns
- A Mixture of Generalized Hyperbolic Distributions
- Simultaneous model-based clustering and visualization in the Fisher discriminative subspace
- Mixtures of Skew-t Factor Analyzers
- Parsimonious Skew Mixture Models for Model-Based Clustering and Classification
- Dimension reduction for model-based clustering
- Mixtures of Common Skew-t Factor Analyzers
- Unsupervised Learning via Mixtures of Skewed Distributions with Hypercube Contours
- clustvarsel: A Package Implementing Variable Selection for Model-based Clustering in R
- Fractionally-Supervised Classification
- Mixtures of Variance-Gamma Distributions
- Graphical tools for model-based mixture discriminant analysis
Cited by in corpus (8)
- Variational Bayes Approximations for Clustering via Mixtures of Normal Inverse Gaussian Distributions
- Mixtures of Generalized Hyperbolic Distributions and Mixtures of Skew-t Distributions for Model-Based Clustering with Incomplete Data
- A Mixture of SDB Skew-t Factor Analyzers
- Hidden Truncation Hyperbolic Distributions, Finite Mixtures Thereof, and Their Application for Clustering
- Flexible Clustering for High-Dimensional Data via Mixtures of Joint Generalized Hyperbolic Models
- Variable selection for mixed data clustering: a model-based approach
- Flexible High-Dimensional Unsupervised Learning with Missing Data
- Clustering Airbnb Reviews