Benchmarking in cluster analysis: A white paper
arXiv:1809.10496 · doi:10.1002/widm.1511
Abstract
Note: A revised version of this is now published. Please cite and read (it's open access): Van Mechelen, I., Boulesteix, A.-L., Dangl, R., Dean, N., Hennig, C., Leisch, F., Steinley, D., Warrens, M. J. (2023). A white paper on good research practices in benchmarking: The case of cluster analysis. WIREs Data Mining and Knowledge Discovery, e1511. https://doi.org/10.1002/widm.1511 To achieve scientific progress in terms of building a cumulative body of knowledge, careful attention to benchmarking is of the utmost importance. This means that proposals of new methods of data pre-processing, new data-analytic techniques, and new methods of output post-processing, should be extensively and carefully compared with existing alternatives, and that existing methods should be subjected to neutral comparison studies. To date, benchmarking and recommendations for benchmarking have been frequently seen in the context of supervised learning. Unfortunately, there has been a dearth of guidelines for benchmarking in an unsupervised setting, with the area of clustering as an important subdomain. To address this problem, discussion is given to the theoretical conceptual underpinnings of benchmarking in the field of cluster analysis by means of simulated as well as empirical data. Subsequently, the practicalities of how to address benchmarking questions in clustering are dealt with, and foundational recommendations are made.
References in corpus (8)
- The importance of transparency and reproducibility in artificial intelligence research
- A Benchmark Study on Time Series Clustering
- Are Cluster Validity Measures (In)valid?
- Benchmark and application of unsupervised classification approaches for univariate data
- A framework for benchmarking clustering algorithms
- Selecting the number of clusters, clustering models, and algorithms. A unifying approach based on the quadratic discriminant score
- HAWKS: Evolving Challenging Benchmark Sets for Cluster Analysis
- Cross-Study Replicability in Cluster Analysis
Cited by in corpus (7)
- Clustering with minimum spanning trees: How good can it be?
- Comparing Clustering Approaches for Smart Meter Time Series: Investigating the Influence of Dataset Properties on Performance
- Normalised clustering accuracy: An asymmetric external cluster validity measure
- On "Confirmatory" Methodological Research in Statistics and Related Fields
- Onset of a conceptual outline map to get a hold on the jungle of cluster analysis
- Benchmarking of Clustering Validity Measures Revisited
- Unbiased mixed variables distance