activity
20172022
most citedNeural Language Modeling with Visual Features

23 citations · 58 across the 6 of their papers we have counts for

collaborators

7 papers

eess.AS2022

End-to-End Multi-Person Audio/Visual Automatic Speech Recognition

Otavio Braga, Takaki Makino, Olivier Siohan +1

Traditionally, audio-visual automatic speech recognition has been studied under the assumption that the speaking face on the visual signal is the face matching the audio. However,…

eess.AS201915 cited

Recurrent Neural Network Transducer for Audio-Visual Speech Recognition

Takaki Makino, Hank Liao, Yannis Assael +4

This work presents a large-scale audio-visual speech recognition system based on a recurrent neural network transducer (RNN-T) architecture. To support the development of such a sy…

eess.AS201913 cited

A comparison of end-to-end models for long-form speech recognition

Chung-Cheng Chiu, Wei Han, Yu Zhang +11

End-to-end automatic speech recognition (ASR) models, including both attention-based models and the recurrent neural network transducer (RNN-T), have shown superior performance com…

cs.CL20196 cited

Adversarial Training for Multilingual Acoustic Modeling

Ke Hu, Hasim Sak, Hank Liao

Multilingual training has been shown to improve acoustic modeling performance by sharing and transferring knowledge in modeling different languages. Knowledge sharing is usually ac…

cs.CL201923 cited

Neural Language Modeling with Visual Features

Antonios Anastasopoulos, Shankar Kumar, Hank Liao

Multimodal language models attempt to incorporate non-linguistic features for the language modeling task. In this work, we extend a standard recurrent neural network (RNN) language…

cs.CV2018

Large-Scale Visual Speech Recognition

Brendan Shillingford, Yannis Assael, Matthew W. Hoffman +12

This work presents a scalable solution to open-vocabulary visual speech recognition. To achieve this, we constructed the largest existing visual speech recognition dataset, consist…