2 citations · 2 across the 2 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2021
VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation
Changhan Wang, Morgane Rivière, Ann Lee +6
We introduce VoxPopuli, a large-scale multilingual corpus providing 100K hours of unlabelled speech data in 23 languages. It is the largest open data to date for unsupervised repre…
cs.CL2020
Population Based Training for Data Augmentation and Regularization in Speech Recognition
Daniel Haziza, Jérémy Rapin, Gabriel Synnaeve
Varying data augmentation policies and regularization over the course of optimization has led to performance improvements over using fixed values. We show that population based tra…