23 citations · 58 across the 6 of their papers we have counts for
4 papers · 1 filter
On Robustness to Missing Video for Audiovisual Speech Recognition
Oscar Chang, Otavio Braga, Hank Liao +2
It has been shown that learning audiovisual features can lead to improved speech recognition performance over audio-only features, especially for noisy speech. However, in many com…
End-to-End Multi-Person Audio/Visual Automatic Speech Recognition
Otavio Braga, Takaki Makino, Olivier Siohan +1
Traditionally, audio-visual automatic speech recognition has been studied under the assumption that the speaking face on the visual signal is the face matching the audio. However,…
Recurrent Neural Network Transducer for Audio-Visual Speech Recognition
Takaki Makino, Hank Liao, Yannis Assael +4
This work presents a large-scale audio-visual speech recognition system based on a recurrent neural network transducer (RNN-T) architecture. To support the development of such a sy…
A comparison of end-to-end models for long-form speech recognition
Chung-Cheng Chiu, Wei Han, Yu Zhang +11
End-to-end automatic speech recognition (ASR) models, including both attention-based models and the recurrent neural network transducer (RNN-T), have shown superior performance com…