most citedLarge-Scale Self- and Semi-Supervised Learning for Speech Translation

7 citations · 7 across the 1 of their papers we have counts for

collaborators

5 papers

cs.CL20217 cited

Large-Scale Self- and Semi-Supervised Learning for Speech Translation

Changhan Wang, Anne Wu, Juan Pino +3

In this paper, we improve speech translation (ST) through effectively leveraging large quantities of unlabeled speech and text data in different and complementary ways. We explore…

cs.CL2021

VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Changhan Wang, Morgane Rivière, Ann Lee +6

We introduce VoxPopuli, a large-scale multilingual corpus providing 100K hours of unlabelled speech data in 23 languages. It is the largest open data to date for unsupervised repre…

cs.CL2020

CoVoST 2 and Massively Multilingual Speech-to-Text Translation

Changhan Wang, Anne Wu, Juan Pino

Speech translation has recently become an increasingly popular topic of research, partly due to the development of benchmark datasets. Nevertheless, current datasets cover a limite…

eess.AS2020

Self-Supervised Representations Improve End-to-End Speech Translation

Anne Wu, Changhan Wang, Juan Pino +1

End-to-end speech-to-text translation can provide a simpler and smaller system but is facing the challenge of data scarcity. Pre-training methods can leverage unlabeled data and ha…

cs.CL2020

CoVoST: A Diverse Multilingual Speech-To-Text Translation Corpus

Changhan Wang, Juan Pino, Anne Wu +1

Spoken language translation has recently witnessed a resurgence in popularity, thanks to the development of end-to-end models and the creation of new corpora, such as Augmented Lib…