68 citations · 227 across the 28 of their papers we have counts for
4 papers · 1 filter
fairseq S^2: A Scalable and Integrable Speech Synthesis Toolkit
Changhan Wang, Wei-Ning Hsu, Yossi Adi +5
This paper presents fairseq S^2, a fairseq extension for speech synthesis. We implement a number of autoregressive (AR) and non-AR text-to-speech models, and their multi-speaker va…
Self-Supervised Representations Improve End-to-End Speech Translation
Anne Wu, Changhan Wang, Juan Pino +1
End-to-end speech-to-text translation can provide a simpler and smaller system but is facing the challenge of data scarcity. Pre-training methods can leverage unlabeled data and ha…
Improving Cross-Lingual Transfer Learning for End-to-End Speech Recognition with Speech Translation
Changhan Wang, Juan Pino, Jiatao Gu
Transfer learning from high-resource languages is known to be an efficient way to improve end-to-end automatic speech recognition (ASR) for low-resource languages. Pre-trained or j…
SkinAugment: Auto-Encoding Speaker Conversions for Automatic Speech Translation
Arya D. McCarthy, Liezl Puzon, Juan Pino
We propose autoencoding speaker conversion for training data augmentation in automatic speech translation. This technique directly transforms an audio sequence, resulting in audio…