5 papers
Internal Evaluation of Density-Based Clusterings with Noise
Anna Beer, Lena Krieger, Pascal Weber +3
Being able to evaluate the quality of a clustering result even in the absence of ground truth cluster labels is fundamental for research in data mining. However, most cluster valid…
AdaBoost is not an Optimal Weak to Strong Learner
Mikael Møller Høgsgaard, Kasper Green Larsen, Martin Ritzert
AdaBoost is a classic boosting algorithm for combining multiple inaccurate classifiers produced by a weak learner, to produce a strong learner with arbitrarily high accuracy when g…
Hierarchical clustering with maximum density paths and mixture models
Martin Ritzert, Polina Turishcheva, Laura Hansel +3
Hierarchical clustering is an effective, interpretable method for analyzing structure in data. It reveals insights at multiple scales without requiring a predefined number of clust…
Boosting, Voting Classifiers and Randomized Sample Compression Schemes
Arthur da Cunha, Kasper Green Larsen, Martin Ritzert
In boosting, we aim to leverage multiple weak learners to produce a strong learner. At the center of this paradigm lies the concept of building the strong learner as a voting class…
MNIST-Nd: a set of naturalistic datasets to benchmark clustering across dimensions
Polina Turishcheva, Laura Hansel, Martin Ritzert +2
Driven by advances in recording technology, large-scale high-dimensional datasets have emerged across many scientific disciplines. Especially in biology, clustering is often used t…