7 papers
SelectTSL: Prompt-Guided Selective Target Sound Localization in Complex Scenarios
Ziyang Jiang, Yu Chen, Zexu Pan +5
Humans can selectively attend to a target sound and estimate its direction in complex scenarios, whereas such selective localization remains challenging for current deep learning-b…
CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction
Jiadong Wang, Ke Zhang, Xinyuan Qian +3
Audio-visual speaker extraction has attracted increasing attention, as it removes the need for pre-registered speech and leverages the visual modality as a complement to audio. Alt…
Voice Conversion Augmentation for Speaker Recognition on Defective Datasets
Ruijie Tao, Zhan Shi, Yidi Jiang +2
Modern speaker recognition system relies on abundant and balanced datasets for classification training. However, diverse defective datasets, such as partially-labelled, small-scale…
Exploring Length Generalization For Transformer-based Speech Enhancement
Qiquan Zhang, Hongxu Zhu, Xinyuan Qian +2
Transformer network architecture has proven effective in speech enhancement. However, as its core module, self-attention suffers from quadratic complexity, making it infeasible for…
Mamba in Speech: Towards an Alternative to Self-Attention
Xiangyu Zhang, Qiquan Zhang, Hexin Liu +6
Transformer and its derivatives have achieved success in diverse tasks across computer vision, natural language processing, and speech processing. To reduce the complexity of compu…
SAV-SE: Scene-aware Audio-Visual Speech Enhancement with Selective State Space Model
Xinyuan Qian, Jiaran Gao, Yaodan Zhang +4
Speech enhancement plays an essential role in various applications, and the integration of visual information has been demonstrated to bring substantial advantages. However, the ma…