2 papers
cs.LG2026
CASA: Classification Augmented with Safety Attention for Robust Multimodal Alignment
Anurag Kumar, Raghuveer Peri, Jon Burnsky +4
Multimodal large-language models (MLLMs) often experience degraded safety alignment when harmful queries exploit cross-modal interactions. Models aligned on text alone show a highe…
eess.AS2025
SEAL: Speaker Error Correction using Acoustic-conditioned Large Language Models
Anurag Kumar, Rohit Paturi, Amber Afshan +1
Speaker Diarization (SD) is a crucial component of modern end-to-end ASR pipelines. Traditional SD systems, which are typically audio-based and operate independently of ASR, often…