15 citations · 17 across the 3 of their papers we have counts for
3 papers
eess.AS2020★ 1 cited
Audio-Visual Decision Fusion for WFST-based and seq2seq Models
Rohith Aralikatti, Sharad Roy, Abhinav Thanda +4
Under noisy conditions, speech recognition systems suffer from high Word Error Rates (WER). In such cases, information from the visual modality comprising the speaker lip movements…
cs.CL2017★ 15 cited
Multi-task Learning Of Deep Neural Networks For Audio Visual Automatic Speech Recognition
Abhinav Thanda, Shankar M Venkatesan
Multi-task learning (MTL) involves the simultaneous training of two or more related tasks over shared representations. In this work, we apply MTL to audio-visual automatic speech r…
cs.CV2016★ 1 cited
Audio Visual Speech Recognition using Deep Recurrent Neural Networks
Abhinav Thanda, Shankar M Venkatesan
In this work, we propose a training algorithm for an audio-visual automatic speech recognition (AV-ASR) system using deep recurrent neural network (RNN).First, we train a deep RNN…