7 citations · 7 across the 1 of their papers we have counts for
5 papers
Large-Scale Self- and Semi-Supervised Learning for Speech Translation
Changhan Wang, Anne Wu, Juan Pino +3
In this paper, we improve speech translation (ST) through effectively leveraging large quantities of unlabeled speech and text data in different and complementary ways. We explore…
VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation
Changhan Wang, Morgane Rivière, Ann Lee +6
We introduce VoxPopuli, a large-scale multilingual corpus providing 100K hours of unlabelled speech data in 23 languages. It is the largest open data to date for unsupervised repre…
CoVoST 2 and Massively Multilingual Speech-to-Text Translation
Changhan Wang, Anne Wu, Juan Pino
Speech translation has recently become an increasingly popular topic of research, partly due to the development of benchmark datasets. Nevertheless, current datasets cover a limite…
Self-Supervised Representations Improve End-to-End Speech Translation
Anne Wu, Changhan Wang, Juan Pino +1
End-to-end speech-to-text translation can provide a simpler and smaller system but is facing the challenge of data scarcity. Pre-training methods can leverage unlabeled data and ha…
CoVoST: A Diverse Multilingual Speech-To-Text Translation Corpus
Changhan Wang, Juan Pino, Anne Wu +1
Spoken language translation has recently witnessed a resurgence in popularity, thanks to the development of end-to-end models and the creation of new corpora, such as Augmented Lib…