24 citations · 26 across the 4 of their papers we have counts for
4 papers
Offline Visual Representation Learning for Embodied Navigation
Karmesh Yadav, Ram Ramrakhya, Arjun Majumdar +5
How should we learn visual representations for embodied agents that must see and move? The status quo is tabula rasa in vivo, i.e. learning visual representations from scratch whil…
On-demand compute reduction with stochastic wav2vec 2.0
Apoorv Vyas, Wei-Ning Hsu, Michael Auli +1
Squeeze and Efficient Wav2vec (SEW) is a recently proposed architecture that squeezes the input to the transformer encoder for compute efficient pre-training and inference with wav…
Unified Speech-Text Pre-training for Speech Translation and Recognition
Yun Tang, Hongyu Gong, Ning Dong +8
We describe a method to jointly pre-train speech and text in an encoder-decoder modeling framework for speech translation and recognition. The proposed method incorporates four sel…
Improved Language Identification Through Cross-Lingual Self-Supervised Learning
Andros Tjandra, Diptanu Gon Choudhury, Frank Zhang +6
Language identification greatly impacts the success of downstream tasks such as automatic speech recognition. Recently, self-supervised speech representations learned by wav2vec 2.…