318 citations · 1.2k across the 25 of their papers we have counts for
9 papers · 1 filter
Reservoir Transformers
Sheng Shen, Alexei Baevski, Ari S. Morcos +3
We demonstrate that transformers obtain impressive performance even when some of the layers are randomly initialized and never updated. Inspired by old and well-established ideas i…
Language Models not just for Pre-training: Fast Online Neural Noisy Channel Modeling
Shruti Bhosale, Kyra Yee, Sergey Edunov +1
Pre-training models on vast quantities of unlabeled data has emerged as an effective approach to improving accuracy on many NLP tasks. On the other hand, traditional machine transl…
A Comparison of Discrete Latent Variable Models for Speech Representation Learning
Henry Zhou, Alexei Baevski, Michael Auli
Neural latent variable models enable the discovery of interesting structure in speech audio data. This paper presents a comparison of two different approaches which are broadly bas…
Self-training and Pre-training are Complementary for Speech Recognition
Qiantong Xu, Alexei Baevski, Tatiana Likhomanenko +5
Self-training and unsupervised pre-training have emerged as effective approaches to improve speech recognition systems using unlabeled data. However, it is not clear whether they l…
Beyond English-Centric Multilingual Machine Translation
Angela Fan, Shruti Bhosale, Holger Schwenk +14
Existing work in translation demonstrated the potential of massively multilingual machine translation by training a single model able to translate between any pair of languages. Ho…
Self-training Improves Pre-training for Natural Language Understanding
Jingfei Du, Edouard Grave, Beliz Gunel +5
Unsupervised pre-training has led to much recent progress in natural language understanding. In this paper, we study self-training as another way to leverage unlabeled data through…