Data clustering: a fundamental method in data science and management
arXiv:2412.18760 · doi:10.1016/j.dsm.2025.08.001
Abstract
This paper explores the critical role of data clustering in data science, emphasizing its methodologies, tools, and diverse applications. Traditional techniques, such as partitional and hierarchical clustering, are analyzed alongside advanced approaches such as data stream, density-based, graph-based, and model-based clustering for handling complex structured datasets. The paper highlights key principles underpinning clustering, outlines widely used tools and frameworks, introduces the workflow of clustering in data science, discusses challenges in practical implementation, and examines various applications of clustering. By focusing on these foundations and applications, the discussion underscores clustering's transformative potential. The paper concludes with insights into future research directions, emphasizing clustering's role in driving innovation and enabling data-driven decision-making.
Data Science and Management (2025)
References in corpus (13)
- Fast unfolding of communities in large networks
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
- From Louvain to Leiden: guaranteeing well-connected communities
- Clustering and Community Detection in Directed Networks: A Survey
- Large Scale Adversarial Representation Learning
- SPICE: Semantic Pseudo-labeling for Image Clustering
- Estimating the Optimal Number of Clusters in Categorical Data Clustering by Silhouette Coefficient
- Self-labelling via simultaneous clustering and representation learning
- DeepVATS: Deep Visual Analytics for Time Series
- DouFu: A Double Fusion Joint Learning Method For Driving Trajectory Representation
- Clustering Method for Time-Series Images Using Quantum-Inspired Computing Technology
- Interpreting the Curse of Dimensionality from Distance Concentration and Manifold Effect
- Analyzing categorical time series with the R package ctsfeatures