102 citations · 112 across the 11 of their papers we have counts for
6 papers · 1 filter
Audio Deep Fake Detection System with Neural Stitching for ADD 2022
Rui Yan, Cheng Wen, Shuran Zhou +3
This paper describes our best system and methodology for ADD 2022: The First Audio Deep Synthesis Detection Challenge\cite{Yi2022ADD}. The very same system was used for both two ro…
Time Domain Adversarial Voice Conversion for ADD 2022
Cheng Wen, Tingwei Guo, Xingjun Tan +5
In this paper, we describe our speech generation system for the first Audio Deep Synthesis Detection Challenge (ADD 2022). Firstly, we build an any-to-many voice conversion (VC) sy…
Semantic Data Augmentation for End-to-End Mandarin Speech Recognition
Jianwei Sun, Zhiyuan Tang, Hengxin Yin +6
End-to-end models have gradually become the preferred option for automatic speech recognition (ASR) applications. During the training of end-to-end ASR, data augmentation is a quit…
DiDiSpeech: A Large Scale Mandarin Speech Corpus
Tingwei Guo, Cheng Wen, Dongwei Jiang +8
This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the…
Transformer based unsupervised pre-training for acoustic representation learning
Ruixiong Zhang, Haiwei Wu, Wubo Li +3
Recently, a variety of acoustic tasks and related applications arised. For many acoustic tasks, the labeled data size may be limited. To handle this problem, we propose an unsuperv…
A Further Study of Unsupervised Pre-training for Transformer Based Speech Recognition
Dongwei Jiang, Wubo Li, Ruixiong Zhang +5
Building a good speech recognition system usually requires large amounts of transcribed data, which is expensive to collect. To tackle this problem, many unsupervised pre-training…