14 citations · 15 across the 2 of their papers we have counts for
2 papers
eess.AS2020★ 1 cited
Audio-Visual Decision Fusion for WFST-based and seq2seq Models
Rohith Aralikatti, Sharad Roy, Abhinav Thanda +4
Under noisy conditions, speech recognition systems suffer from high Word Error Rates (WER). In such cases, information from the visual modality comprising the speaker lip movements…
cs.CV2019★ 14 cited
LipReading with 3D-2D-CNN BLSTM-HMM and word-CTC models
Dilip Kumar Margam, Rohith Aralikatti, Tanay Sharma +4
In recent years, deep learning based machine lipreading has gained prominence. To this end, several architectures such as LipNet, LCANet and others have been proposed which perform…