3 papers
cs.SD2026
Soft Posterior Speaker Injection for Multi-Talker Speech Recognition
Jian Zhu, Cheng Luo
Multi-talker automatic speech recognition (MT-ASR) remains challenging in the presence of overlapping speech. Hard segmentation introduces irreversible errors, whereas serialized o…
cs.CV2025
MoEGCL: Mixture of Ego-Graphs Contrastive Representation Learning for Multi-View Clustering
Jian Zhu, Xin Zou, Jun Sun +7
In recent years, the advancement of Graph Neural Networks (GNNs) has significantly propelled progress in Multi-View Clustering (MVC). However, existing methods face the problem of…
cs.CV2024
CLIP Multi-modal Hashing for Multimedia Retrieval
Jian Zhu, Mingkai Sheng, Zhangmin Huang +5
Multi-modal hashing methods are widely used in multimedia retrieval, which can fuse multi-source data to generate binary hash code. However, the individual backbone networks have l…