3 papers
cs.SD2025
Emotional Styles Hide in Deep Speaker Embeddings: Disentangle Deep Speaker Embeddings for Speaker Clustering
Chaohao Lin, Xu Zheng, Kaida Wu +2
Speaker clustering is the task of identifying the unique speakers in a set of audio recordings (each belonging to exactly one speaker) without knowing who and how many speakers are…
cs.CV2025
Parameter-Efficient CLIP Adaptation for 3D Understanding via Unified Tokenization
Guofeng Mei, Qinfeng Xiao, Bin Ren +7
Vision-language models, such as CLIP, encode rich semantic knowledge through large-scale image-text pretraining. Reusing these models for 3D understanding is highly desirable, beca…
cs.CV2025
Efficient Degradation-agnostic Image Restoration via Channel-Wise Functional Decomposition and Manifold Regularization
Bin Ren, Yawei Li, Xu Zheng +6
Degradation-agnostic image restoration aims to handle diverse corruptions with one unified model, but faces fundamental challenges in balancing efficiency and performance across di…