23 citations · 58 across the 6 of their papers we have counts for
7 papers
End-to-End Multi-Person Audio/Visual Automatic Speech Recognition
Otavio Braga, Takaki Makino, Olivier Siohan +1
Traditionally, audio-visual automatic speech recognition has been studied under the assumption that the speaking face on the visual signal is the face matching the audio. However,…
Recurrent Neural Network Transducer for Audio-Visual Speech Recognition
Takaki Makino, Hank Liao, Yannis Assael +4
This work presents a large-scale audio-visual speech recognition system based on a recurrent neural network transducer (RNN-T) architecture. To support the development of such a sy…
A comparison of end-to-end models for long-form speech recognition
Chung-Cheng Chiu, Wei Han, Yu Zhang +11
End-to-end automatic speech recognition (ASR) models, including both attention-based models and the recurrent neural network transducer (RNN-T), have shown superior performance com…
Adversarial Training for Multilingual Acoustic Modeling
Ke Hu, Hasim Sak, Hank Liao
Multilingual training has been shown to improve acoustic modeling performance by sharing and transferring knowledge in modeling different languages. Knowledge sharing is usually ac…
Neural Language Modeling with Visual Features
Antonios Anastasopoulos, Shankar Kumar, Hank Liao
Multimodal language models attempt to incorporate non-linguistic features for the language modeling task. In this work, we extend a standard recurrent neural network (RNN) language…
Large-Scale Visual Speech Recognition
Brendan Shillingford, Yannis Assael, Matthew W. Hoffman +12
This work presents a scalable solution to open-vocabulary visual speech recognition. To achieve this, we constructed the largest existing visual speech recognition dataset, consist…