-POD: A Method for -Means Clustering of Missing Data
arXiv:1411.7013 · doi:10.1080/00031305.2015.1086685
Abstract
The -means algorithm is often used in clustering applications but its usage requires a complete data matrix. Missing data, however, is common in many applications. Mainstream approaches to clustering missing data reduce the missing data problem to a complete data formulation through either deletion or imputation but these solutions may incur significant costs. Our -POD method presents a simple extension of -means clustering for missing data that works even when the missingness mechanism is unknown, when external information is unavailable, and when there is significant missingness in the data.
26 pages, 7 tables
References in corpus (3)
Cited by in corpus (9)
- An efficient -means-type algorithm for clustering datasets with incomplete records
- Fast model-based clustering of partial records
- Simple and Scalable Sparse k-means Clustering via Feature Ranking
- Estimating Heterogeneous Causal Effects of High-Dimensional Treatments: Application to Conjoint Analysis
- Flexible High-Dimensional Unsupervised Learning with Missing Data
- Clustering Data with Nonignorable Missingness using Semi-Parametric Mixture Models
- Optimal Clustering with Missing Values
- Clustering multilayer graphs with missing nodes
- Clustering with missing data: which imputation model for which cluster analysis method?