papers

Publications (39)

cs.LG2020

Self-training and Pre-training are Complementary for Speech Recognition

Qiantong Xu, Alexei Baevski, Tatiana Likhomanenko +5

Self-training and unsupervised pre-training have emerged as effective approaches to improve speech recognition systems using unlabeled data. However, it is not clear whether they l…

cs.SD2021

Robust wav2vec 2.0: Analyzing Domain Shift in Self-Supervised Pre-Training

Wei-Ning Hsu, Anuroop Sriram, Alexei Baevski +8

Self-supervised learning of speech representations has been a very active research area but most work is focused on a single domain such as read audio books for which there exist l…

cs.CL2022

Towards End-to-end Unsupervised Speech Recognition

Alexander H. Liu, Wei-Ning Hsu, Michael Auli +1

Unsupervised speech recognition has shown great potential to make Automatic Speech Recognition (ASR) systems accessible to every language. However, existing methods still heavily r…

cs.CL2023

Toward Joint Language Modeling for Speech Units and Text

Ju-Chieh Chou, Chung-Ming Chien, Wei-Ning Hsu +5

Speech and text are two major forms of human language. The research community has been focusing on mapping speech to text or vice versa for many years. However, in the field of lan…

cs.CL2019

Adaptive Input Representations for Neural Language Modeling

Alexei Baevski, Michael Auli

We introduce adaptive input representations for neural language modeling which extend the adaptive softmax of Grave et al. (2017) to input representations of variable capacity. The…

cs.LG2022

On-demand compute reduction with stochastic wav2vec 2.0

Apoorv Vyas, Wei-Ning Hsu, Michael Auli +1

Squeeze and Efficient Wav2vec (SEW) is a recently proposed architecture that squeezes the input to the transformer encoder for compute efficient pre-training and inference with wav…