12 papers
Imputation Meets Clustering: Exploiting Latent Subgroup Structure for Missing Data Recovery
Chuyao Zhang, E Li, Taochen Chen +5
Missing data is prevalent in practical applications, making effective imputation an essential preprocessing step for downstream analysis. Real-world datasets often exhibit complex…
Bridging the Semantic Gap for Categorical Data Clustering via Large Language Models
Zihua Yang, Xin Liao, Yiqun Zhang +1
Qualitative data are widespread in domains such as healthcare, marketing, and bioinformatics, where clustering offers a fundamental tool for pattern discovery. A core difficulty of…
Beyond Statistical Co-occurrence: Unlocking Intrinsic Semantics for Tabular Data Clustering
Mingjie Zhao, Yunfan Zhang, Yiqun Zhang +1
Deep Clustering (DC) has emerged as a powerful tool for tabular data analysis in real-world domains like finance and healthcare. However, most existing methods rely on data-level s…
CADM: Cluster-customized Adaptive Distance Metric for Categorical Data Clustering
Taixi Chen, Yiu-ming Cheung, Yiqun Zhang
An appropriate distance metric is crucial for categorical data clustering, as the distance between categorical data cannot be directly calculated. However, the distances between at…
Robust Categorical Data Clustering Guided by Multi-Granular Competitive Learning
Shenghong Cai, Yiqun Zhang, Xiaopeng Luo +3
Data set composed of categorical features is very common in big data analysis tasks. Since categorical features are usually with a limited number of qualitative possible values, th…
Evo-TFS: Evolutionary Time-Frequency Domain-Based Synthetic Minority Oversampling Approach to Imbalanced Time Series Classification
Wenbin Pei, Ruohao Dai, Bing Xue +3
Time series classification is a fundamental machine learning task with broad real-world applications. Although many deep learning methods have proven effective in learning time-ser…