15 citations · 15 across the 2 of their papers we have counts for
2 papers
eess.AS2022
End-to-End Multi-Person Audio/Visual Automatic Speech Recognition
Otavio Braga, Takaki Makino, Olivier Siohan +1
Traditionally, audio-visual automatic speech recognition has been studied under the assumption that the speaking face on the visual signal is the face matching the audio. However,…
eess.AS2019★ 15 cited
Recurrent Neural Network Transducer for Audio-Visual Speech Recognition
Takaki Makino, Hank Liao, Yannis Assael +4
This work presents a large-scale audio-visual speech recognition system based on a recurrent neural network transducer (RNN-T) architecture. To support the development of such a sy…