activity
20182025
most citedLipReading with 3D-2D-CNN BLSTM-HMM and word-CTC models

14 citations · 16 across the 6 of their papers we have counts for

collaborators

7 papers

cs.CV2025

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings

Aakriti Agrawal, Gouthaman KV, Rohith Aralikatti +6

Hallucinations in Large Vision-Language Models (LVLMs) remain a persistent challenge, often stemming from inadequate integration of visual information during multimodal reasoning.…

cs.CL2025

Uncertainty-Aware Answer Selection for Improved Reasoning in Multi-LLM Systems

Aakriti Agrawal, Rohith Aralikatti, Anirudh Satheesh +3

Large Language Models (LLMs) have demonstrated exceptional capabilities, yet selecting the most reliable response from multiple LLMs remains a challenge, particularly in resource-c…

eess.AS2022

Reverberation as Supervision for Speech Separation

Rohith Aralikatti, Christoph Boeddeker, Gordon Wichern +2

This paper proposes reverberation as supervision (RAS), a novel unsupervised loss function for single-channel reverberant speech separation. Prior methods for unsupervised separati…

eess.AS20211 cited

Improving Reverberant Speech Separation with Multi-stage Training and Curriculum Learning

Rohith Aralikatti, Anton Ratnarajah, Zhenyu Tang +1

We present a novel approach that improves the performance of reverberant speech separation. Our approach is based on an accurate geometric acoustic simulator (GAS) which generates…

eess.AS20201 cited

Audio-Visual Decision Fusion for WFST-based and seq2seq Models

Rohith Aralikatti, Sharad Roy, Abhinav Thanda +4

Under noisy conditions, speech recognition systems suffer from high Word Error Rates (WER). In such cases, information from the visual modality comprising the speaker lip movements…

cs.CV201914 cited

LipReading with 3D-2D-CNN BLSTM-HMM and word-CTC models

Dilip Kumar Margam, Rohith Aralikatti, Tanay Sharma +4

In recent years, deep learning based machine lipreading has gained prominence. To this end, several architectures such as LipNet, LCANet and others have been proposed which perform…