activity
20242026
collaborators

8 papers

cs.LG2026

Poisson Subspace Clustering: Focusing on the Essentials in Count Data

Collin Leiber, Kai Puolamäki, Heikki Mannila

Count data represented as a matrix of non-negative integer values, such as contingency tables, are prevalent across diverse domains. When clustering such data sets, specific method…

cs.LG2026

Khatri-Rao Clustering for Data Summarization

Martino Ciaperoni, Collin Leiber, Aristides Gionis +1

As datasets continue to grow in size and complexity, finding succinct yet accurate data summaries poses a key challenge. Centroid-based clustering, a widely adopted approach to add…

cs.LG2025

An Introductory Survey to Autoencoder-based Deep Clustering -- Sandboxes for Combining Clustering with Deep Learning

Collin Leiber, Lukas Miklautz, Claudia Plant +1

Autoencoders offer a general way of learning low-dimensional, non-linear representations from data without labels. This is achieved without making any particular assumptions about…

cs.LG2025

Automatic Parameter Selection for Non-Redundant Clustering

Collin Leiber, Dominik Mautz, Claudia Plant +1

High-dimensional datasets often contain multiple meaningful clusterings in different subspaces. For example, objects can be clustered either by color, weight, or size, revealing di…

cs.LG2025

Extension of the Dip-test Repertoire -- Efficient and Differentiable p-value Calculation for Clustering

Lena G. M. Bauer, Collin Leiber, Christian Böhm +1

Over the last decade, the Dip-test of unimodality has gained increasing interest in the data mining community as it is a parameter-free statistical test that reliably rates the mod…

cs.LG2025

Breaking the Reclustering Barrier in Centroid-based Deep Clustering

Lukas Miklautz, Timo Klein, Kevin Sidak +5

This work investigates an important phenomenon in centroid-based deep clustering (DC) algorithms: Performance quickly saturates after a period of rapid early gains. Practitioners c…