paper

-POD: A Method for -Means Clustering of Missing Data

arXiv:1411.7013 · doi:10.1080/00031305.2015.1086685

Abstract

The -means algorithm is often used in clustering applications but its usage requires a complete data matrix. Missing data, however, is common in many applications. Mainstream approaches to clustering missing data reduce the missing data problem to a complete data formulation through either deletion or imputation but these solutions may incur significant costs. Our -POD method presents a simple extension of -means clustering for missing data that works even when the missingness mechanism is unknown, when external information is unavailable, and when there is significant missingness in the data.

26 pages, 7 tables

References in corpus (3)

Cited by in corpus (9)