14 citations · 16 across the 6 of their papers we have counts for
7 papers
Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings
Aakriti Agrawal, Gouthaman KV, Rohith Aralikatti +6
Hallucinations in Large Vision-Language Models (LVLMs) remain a persistent challenge, often stemming from inadequate integration of visual information during multimodal reasoning.…
Uncertainty-Aware Answer Selection for Improved Reasoning in Multi-LLM Systems
Aakriti Agrawal, Rohith Aralikatti, Anirudh Satheesh +3
Large Language Models (LLMs) have demonstrated exceptional capabilities, yet selecting the most reliable response from multiple LLMs remains a challenge, particularly in resource-c…
Reverberation as Supervision for Speech Separation
Rohith Aralikatti, Christoph Boeddeker, Gordon Wichern +2
This paper proposes reverberation as supervision (RAS), a novel unsupervised loss function for single-channel reverberant speech separation. Prior methods for unsupervised separati…
Improving Reverberant Speech Separation with Multi-stage Training and Curriculum Learning
Rohith Aralikatti, Anton Ratnarajah, Zhenyu Tang +1
We present a novel approach that improves the performance of reverberant speech separation. Our approach is based on an accurate geometric acoustic simulator (GAS) which generates…
Audio-Visual Decision Fusion for WFST-based and seq2seq Models
Rohith Aralikatti, Sharad Roy, Abhinav Thanda +4
Under noisy conditions, speech recognition systems suffer from high Word Error Rates (WER). In such cases, information from the visual modality comprising the speaker lip movements…
LipReading with 3D-2D-CNN BLSTM-HMM and word-CTC models
Dilip Kumar Margam, Rohith Aralikatti, Tanay Sharma +4
In recent years, deep learning based machine lipreading has gained prominence. To this end, several architectures such as LipNet, LCANet and others have been proposed which perform…