activity
20242026
collaborators

12 papers

cs.LG2026

Imputation Meets Clustering: Exploiting Latent Subgroup Structure for Missing Data Recovery

Chuyao Zhang, E Li, Taochen Chen +5

Missing data is prevalent in practical applications, making effective imputation an essential preprocessing step for downstream analysis. Real-world datasets often exhibit complex…

cs.LG2026

Bridging the Semantic Gap for Categorical Data Clustering via Large Language Models

Zihua Yang, Xin Liao, Yiqun Zhang +1

Qualitative data are widespread in domains such as healthcare, marketing, and bioinformatics, where clustering offers a fundamental tool for pattern discovery. A core difficulty of…

cs.AI2026

Beyond Statistical Co-occurrence: Unlocking Intrinsic Semantics for Tabular Data Clustering

Mingjie Zhao, Yunfan Zhang, Yiqun Zhang +1

Deep Clustering (DC) has emerged as a powerful tool for tabular data analysis in real-world domains like finance and healthcare. However, most existing methods rely on data-level s…

cs.LG2026

CADM: Cluster-customized Adaptive Distance Metric for Categorical Data Clustering

Taixi Chen, Yiu-ming Cheung, Yiqun Zhang

An appropriate distance metric is crucial for categorical data clustering, as the distance between categorical data cannot be directly calculated. However, the distances between at…

cs.LG2026

Robust Categorical Data Clustering Guided by Multi-Granular Competitive Learning

Shenghong Cai, Yiqun Zhang, Xiaopeng Luo +3

Data set composed of categorical features is very common in big data analysis tasks. Since categorical features are usually with a limited number of qualitative possible values, th…

cs.LG2026

Evo-TFS: Evolutionary Time-Frequency Domain-Based Synthetic Minority Oversampling Approach to Imbalanced Time Series Classification

Wenbin Pei, Ruohao Dai, Bing Xue +3

Time series classification is a fundamental machine learning task with broad real-world applications. Although many deep learning methods have proven effective in learning time-ser…