318 citations · 876 across the 8 of their papers we have counts for
16 papers · 1 filter
Simple and Effective Zero-shot Cross-lingual Phoneme Recognition
Qiantong Xu, Alexei Baevski, Michael Auli
Recent progress in self-training, self-supervised pretraining and unsupervised learning enabled well performing speech recognition systems without any labeled data. However, in man…
Large-Scale Self- and Semi-Supervised Learning for Speech Translation
Changhan Wang, Anne Wu, Juan Pino +3
In this paper, we improve speech translation (ST) through effectively leveraging large quantities of unlabeled speech and text data in different and complementary ways. We explore…
Generative Spoken Language Modeling from Raw Audio
Kushal Lakhotia, Evgeny Kharitonov, Wei-Ning Hsu +8
We introduce Generative Spoken Language Modeling, the task of learning the acoustic and linguistic characteristics of a language from raw audio (no text, no labels), and a set of m…
Reservoir Transformers
Sheng Shen, Alexei Baevski, Ari S. Morcos +3
We demonstrate that transformers obtain impressive performance even when some of the layers are randomly initialized and never updated. Inspired by old and well-established ideas i…
The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling
Tu Anh Nguyen, Maureen de Seyssel, Patricia Rozé +5
We introduce a new unsupervised task, spoken language modeling: the learning of linguistic representations from raw audio signals without any labels, along with the Zero Resource S…
Unsupervised Cross-lingual Representation Learning for Speech Recognition
Alexis Conneau, Alexei Baevski, Ronan Collobert +2
This paper presents XLSR which learns cross-lingual speech representations by pretraining a single model from the raw waveform of speech in multiple languages. We build on wav2vec…