14 citations · 15 across the 2 of their papers we have counts for
3 papers
Audio-Visual Decision Fusion for WFST-based and seq2seq Models
Rohith Aralikatti, Sharad Roy, Abhinav Thanda +4
Under noisy conditions, speech recognition systems suffer from high Word Error Rates (WER). In such cases, information from the visual modality comprising the speaker lip movements…
LipReading with 3D-2D-CNN BLSTM-HMM and word-CTC models
Dilip Kumar Margam, Rohith Aralikatti, Tanay Sharma +4
In recent years, deep learning based machine lipreading has gained prominence. To this end, several architectures such as LipNet, LCANet and others have been proposed which perform…
Global SNR Estimation of Speech Signals using Entropy and Uncertainty Estimates from Dropout Networks
Rohith Aralikatti, Dilip Margam, Tanay Sharma +2
This paper demonstrates two novel methods to estimate the global SNR of speech signals. In both methods, Deep Neural Network-Hidden Markov Model (DNN-HMM) acoustic model used in sp…