activity
20192022
most citedRecurrent Neural Network Transducer for Audio-Visual Speech Recognition

15 citations · 16 across the 6 of their papers we have counts for

collaborators

7 papers

eess.AS2022

A Closer Look at Audio-Visual Multi-Person Speech Recognition and Active Speaker Selection

Otavio Braga, Olivier Siohan

Audio-visual automatic speech recognition is a promising approach to robust ASR under noisy conditions. However, up until recently it had been traditionally studied in isolation as…

eess.AS2022

End-to-End Multi-Person Audio/Visual Automatic Speech Recognition

Otavio Braga, Takaki Makino, Olivier Siohan +1

Traditionally, audio-visual automatic speech recognition has been studied under the assumption that the speaking face on the visual signal is the face matching the audio. However,…

eess.AS2022

Best of Both Worlds: Multi-task Audio-Visual Automatic Speech Recognition and Active Speaker Detection

Otavio Braga, Olivier Siohan

Under noisy conditions, automatic speech recognition (ASR) can greatly benefit from the addition of visual signals coming from a video of the speaker's face. However, when multiple…

cs.SD20221 cited

End-to-end multi-talker audio-visual ASR using an active speaker attention module

Richard Rose, Olivier Siohan

This paper presents a new approach for end-to-end audio-visual multi-talker speech recognition. The approach, referred to here as the visual context attention model (VCAM), is impo…

cs.CV2021

Audio-Visual Speech Recognition is Worth 32328 Voxels

Dmitriy Serdyuk, Otavio Braga, Olivier Siohan

Audio-visual automatic speech recognition (AV-ASR) introduces the video modality into the speech recognition process, often by relying on information conveyed by the motion of the…

cs.CL2021

Bridging the gap between streaming and non-streaming ASR systems bydistilling ensembles of CTC and RNN-T models

Thibault Doutre, Wei Han, Chung-Cheng Chiu +3

Streaming end-to-end automatic speech recognition (ASR) systems are widely used in everyday applications that require transcribing speech to text in real-time. Their minimal latenc…