paper

Handling mild outliers and unobserved values in compositional datasets using finite mixtures of mean-parametrised Dirichlet models

arXiv:2608.23900

Abstract

Heterogeneous compositional data may be simultaneously affected by missing values and atypical points, posing challenges for both clustering and outlier detection. We develop a mixture model for incomplete compositional data under Huber's contamination model, with contamination defined directly on the simplex and on observations that may be missing at random. The model provides a principled representation of outliers and allows the distribution of missing parts to be derived while accounting for contamination. We establish that maximisation of the observed log-likelihood constructed from contaminated, mean-parametrised Dirichlet densities is a convex optimisation problem. We then develop a tailored expectation-maximisation. The E-step incorporates the moments from the distribution of the missing parts of the data. Although the resulting parameter estimates are not available in closed form, the maximisation step admits tractable element-wise iterative updates. Numerical experiments demonstrate the performance of the proposed approach under varying percentage of missingness and contamination, and different sample size. An application to the American Time Use Survey identifies two interpretable clusters corresponding to work-intensive and sociable recreational days, while revealing atypical time-use compositions. In contrast, a conventional Dirichlet mixture model identifies four clusters, reflecting the influence of outliers and an artificial splitting of one cluster.

Handling mild outliers and unobserved values in compositional datasets using finite mixtures of mean-parametrised Dirichlet models · wovepaper