activity
20162023
most citedEnglish Broadcast News Speech Recognition by Humans and Machines

12 citations · 32 across the 27 of their papers we have counts for

collaborators
Showing 2021Show all

13 papers · 1 filter

cs.LG2021★ 1 cited

Loss Landscape Dependent Self-Adjusting Learning Rates in Decentralized Stochastic Gradient Descent

Wei Zhang, Mingrui Liu, Yu Feng +3

Distributed Deep Learning (DDL) is essential for large-scale Deep Learning (DL) training. Synchronous Stochastic Gradient Descent (SSGD) 1 is the de facto DDL optimization method.…

cs.CV2021

Everything at Once -- Multi-modal Fusion Transformer for Video Retrieval

Nina Shvetsova, Brian Chen, Andrew Rouditchenko +6

Multi-modal learning from video data has seen increased attention recently as it allows to train semantically meaningful embeddings without human annotation enabling tasks like zer…

cs.CL2021

Cascaded Multilingual Audio-Visual Learning from Videos

Andrew Rouditchenko, Angie Boggust, David Harwath +8

In this paper, we explore self-supervised audio-visual models that learn from instructional videos. Prior work has shown that these models can relate spoken words and sounds to vis…

cs.CL2021★ 1 cited

Asynchronous Decentralized Distributed Training of Acoustic Models

Xiaodong Cui, Wei Zhang, Abdullah Kayi +5

Large-scale distributed training of deep acoustic models plays an important role in today's high-performance automatic speech recognition (ASR). In this paper we investigate a vari…

cs.CL2021

4-bit Quantization of LSTM-based Speech Recognition Models

Andrea Fasoli, Chia-Yu Chen, Mauricio Serrano +9

We investigate the impact of aggressive low-precision representations of weights and activations in two families of large LSTM-based architectures for Automatic Speech Recognition…

cs.CL2021

Reducing Exposure Bias in Training Recurrent Neural Network Transducers

Xiaodong Cui, Brian Kingsbury, George Saon +2

When recurrent neural network transducers (RNNTs) are trained using the typical maximum likelihood criterion, the prediction network is trained only on ground truth label sequences…