activity
20242026
collaborators

5 papers

cs.LG2026

CASA: Classification Augmented with Safety Attention for Robust Multimodal Alignment

Anurag Kumar, Raghuveer Peri, Jon Burnsky +4

Multimodal large-language models (MLLMs) often experience degraded safety alignment when harmful queries exploit cross-modal interactions. Models aligned on text alone show a highe…

cs.CL2025

SpeechVerse: A Large-scale Generalizable Audio Language Model

Nilaksh Das, Saket Dingliwal, Srikanth Ronanki +14

Large language models (LLMs) have shown incredible proficiency in performing tasks that require semantic understanding of natural language instructions. Recently, many works have f…

eess.AS2025

SEAL: Speaker Error Correction using Acoustic-conditioned Large Language Models

Anurag Kumar, Rohit Paturi, Amber Afshan +1

Speaker Diarization (SD) is a crucial component of modern end-to-end ASR pipelines. Traditional SD systems, which are typically audio-based and operate independently of ASR, often…

eess.AS2024

Speakers Unembedded: Embedding-free Approach to Long-form Neural Diarization

Xiang Li, Vivek Govindan, Rohit Paturi +1

End-to-end neural diarization (EEND) models offer significant improvements over traditional embedding-based Speaker Diarization (SD) approaches but falls short on generalizing to l…

eess.AS2024

AG-LSEC: Audio Grounded Lexical Speaker Error Correction

Rohit Paturi, Xiang Li, Sundararajan Srinivasan

Speaker Diarization (SD) systems are typically audio-based and operate independently of the ASR system in traditional speech transcription pipelines and can have speaker errors due…