20 citations · 133 across the 27 of their papers we have counts for
9 papers · 1 filter
Audio-Visual Active Speaker Extraction for Sparsely Overlapped Multi-talker Speech
Junjie Li, Ruijie Tao, Zexu Pan +3
Target speaker extraction aims to extract the speech of a specific speaker from a multi-talker mixture as specified by an auxiliary reference. Most studies focus on the scenario wh…
USED: Universal Speaker Extraction and Diarization
Junyi Ao, Mehmet Sinan Yıldırım, Ruijie Tao +4
Speaker extraction and diarization are two enabling techniques for real-world speech applications. Speaker extraction aims to extract a target speaker's voice from a speech mixture…
Spiking-LEAF: A Learnable Auditory front-end for Spiking Neural Networks
Zeyang Song, Jibin Wu, Malu Zhang +2
Brain-inspired spiking neural networks (SNNs) have demonstrated great potential for temporal signal processing. However, their performance in speech processing remains limited due…
Betray Oneself: A Novel Audio DeepFake Detection Model via Mono-to-Stereo Conversion
Rui Liu, Jinhua Zhang, Guanglai Gao +1
Audio Deepfake Detection (ADD) aims to detect the fake audio generated by text-to-speech (TTS), voice conversion (VC) and replay, etc., which is an emerging topic. Traditionally we…
ADD 2023: the Second Audio Deepfake Detection Challenge
Jiangyan Yi, Jianhua Tao, Ruibo Fu +15
Audio deepfake detection is an emerging topic in the artificial intelligence community. The second Audio Deepfake Detection Challenge (ADD 2023) aims to spur researchers around the…
Ripple sparse self-attention for monaural speech enhancement
Qiquan Zhang, Hongxu Zhu, Qi Song +3
The use of Transformer represents a recent success in speech enhancement. However, as its core component, self-attention suffers from quadratic complexity, which is computationally…