activity
20182021
most citedPay Less Attention with Lightweight and Dynamic Convolutions

318 citations · 876 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

16 papers · 1 filter

cs.CL202110 cited

Simple and Effective Zero-shot Cross-lingual Phoneme Recognition

Qiantong Xu, Alexei Baevski, Michael Auli

Recent progress in self-training, self-supervised pretraining and unsupervised learning enabled well performing speech recognition systems without any labeled data. However, in man…

cs.CL20217 cited

Large-Scale Self- and Semi-Supervised Learning for Speech Translation

Changhan Wang, Anne Wu, Juan Pino +3

In this paper, we improve speech translation (ST) through effectively leveraging large quantities of unlabeled speech and text data in different and complementary ways. We explore…

cs.CL2021

Generative Spoken Language Modeling from Raw Audio

Kushal Lakhotia, Evgeny Kharitonov, Wei-Ning Hsu +8

We introduce Generative Spoken Language Modeling, the task of learning the acoustic and linguistic characteristics of a language from raw audio (no text, no labels), and a set of m…

cs.CL2020

Reservoir Transformers

Sheng Shen, Alexei Baevski, Ari S. Morcos +3

We demonstrate that transformers obtain impressive performance even when some of the layers are randomly initialized and never updated. Inspired by old and well-established ideas i…

cs.CL202042 cited

The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling

Tu Anh Nguyen, Maureen de Seyssel, Patricia Rozé +5

We introduce a new unsupervised task, spoken language modeling: the learning of linguistic representations from raw audio signals without any labels, along with the Zero Resource S…

cs.CL2020

Unsupervised Cross-lingual Representation Learning for Speech Recognition

Alexis Conneau, Alexei Baevski, Ronan Collobert +2

This paper presents XLSR which learns cross-lingual speech representations by pretraining a single model from the raw waveform of speech in multiple languages. We build on wav2vec…