activity
20162023
most citedPhoenix: Democratizing ChatGPT across Languages

20 citations · 133 across the 27 of their papers we have counts for

collaborators
Showing cs.SDShow all

9 papers · 1 filter

cs.SD2023

Audio-Visual Active Speaker Extraction for Sparsely Overlapped Multi-talker Speech

Junjie Li, Ruijie Tao, Zexu Pan +3

Target speaker extraction aims to extract the speech of a specific speaker from a multi-talker mixture as specified by an auxiliary reference. Most studies focus on the scenario wh…

cs.SD2023

USED: Universal Speaker Extraction and Diarization

Junyi Ao, Mehmet Sinan Yıldırım, Ruijie Tao +4

Speaker extraction and diarization are two enabling techniques for real-world speech applications. Speaker extraction aims to extract a target speaker's voice from a speech mixture…

cs.SD2023

Spiking-LEAF: A Learnable Auditory front-end for Spiking Neural Networks

Zeyang Song, Jibin Wu, Malu Zhang +2

Brain-inspired spiking neural networks (SNNs) have demonstrated great potential for temporal signal processing. However, their performance in speech processing remains limited due…

cs.SD20231 cited

Betray Oneself: A Novel Audio DeepFake Detection Model via Mono-to-Stereo Conversion

Rui Liu, Jinhua Zhang, Guanglai Gao +1

Audio Deepfake Detection (ADD) aims to detect the fake audio generated by text-to-speech (TTS), voice conversion (VC) and replay, etc., which is an emerging topic. Traditionally we…

cs.SD202320 cited

ADD 2023: the Second Audio Deepfake Detection Challenge

Jiangyan Yi, Jianhua Tao, Ruibo Fu +15

Audio deepfake detection is an emerging topic in the artificial intelligence community. The second Audio Deepfake Detection Challenge (ADD 2023) aims to spur researchers around the…

cs.SD2023

Ripple sparse self-attention for monaural speech enhancement

Qiquan Zhang, Hongxu Zhu, Qi Song +3

The use of Transformer represents a recent success in speech enhancement. However, as its core component, self-attention suffers from quadratic complexity, which is computationally…