paper

EPTAS for -means Clustering of Affine Subspaces

arXiv:2010.09580

Abstract

We consider a generalization of the fundamental -means clustering for data with incomplete or corrupted entries. When data objects are represented by points in , a data point is said to be incomplete when some of its entries are missing or unspecified. An incomplete data point with at most unspecified entries corresponds to an axis-parallel affine subspace of dimension at most , called a -point. Thus we seek a partition of input -points into clusters minimizing the -means objective. For , when all coordinates of each point are specified, this is the usual -means clustering. We give an algorithm that finds an -approximate solution in time for some function of , and only.

To be published in Symposium on Discrete Algorithms (SODA) 2021