activity
20162023
most citedSelf-Attention Transducers for End-to-End Speech Recognition

85 citations · 369 across the 37 of their papers we have counts for

collaborators
Showing cs.SDShow all

19 papers · 1 filter

cs.SD202325 cited

Audio Deepfake Detection: A Survey

Jiangyan Yi, Chenglong Wang, Jianhua Tao +3

Audio deepfake detection is an emerging active topic. A growing number of literatures have aimed to study deepfake detection algorithms and achieved effective performance, the prob…

cs.SD20234 cited

Do You Remember? Overcoming Catastrophic Forgetting for Fake Audio Detection

Xiaohui Zhang, Jiangyan Yi, Jianhua Tao +2

Current fake audio detection algorithms have achieved promising performances on most datasets. However, their performance may be significantly degraded when dealing with audio of a…

cs.SD2023

Spatial Reconstructed Local Attention Res2Net with F0 Subband for Fake Speech Detection

Cunhang Fan, Jun Xue, Jianhua Tao +4

The rhythm of bonafide speech is often difficult to replicate, which causes that the fundamental frequency (F0) of synthetic speech is significantly different from that of real spe…

cs.SD2023

TST: Time-Sparse Transducer for Automatic Speech Recognition

Xiaohui Zhang, Mangui Liang, Zhengkun Tian +2

End-to-end model, especially Recurrent Neural Network Transducer (RNN-T), has achieved great success in speech recognition. However, transducer requires a great memory footprint an…

cs.SD20231 cited

Boosting Fast and High-Quality Speech Synthesis with Linear Diffusion

Haogeng Liu, Tao Wang, Jie Cao +2

Denoising Diffusion Probabilistic Models have shown extraordinary ability on various generative tasks. However, their slow inference speed renders them impractical in speech synthe…

cs.SD20235 cited

Low-rank Adaptation Method for Wav2vec2-based Fake Audio Detection

Chenglong Wang, Jiangyan Yi, Xiaohui Zhang +3

Self-supervised speech models are a rapidly developing research topic in fake audio detection. Many pre-trained models can serve as feature extractors, learning richer and higher-l…