5 citations · 9 across the 51 of their papers we have counts for
52 papers
Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models
Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh +1
Audio language models are designed to understand speech, yet it remains unclear whether they capture how something is said beyond what is said. We present a mechanistic analysis of…
Encoder-Decoder Manifold Alignment for Idempotent Generation
Dareen Alharthi, Abdul Waheed, Bhiksha Raj
Recently, several learning paradigms have been introduced to enforce idempotency in generative models. The goal is to ensure that repeated application of a model leaves samples unc…
RIVET: Robust Idempotent Voice Attribute Editing
Dareen Alharthi, Bhuvan Koduru, Rita Singh +1
Voice attribute editing models modify characteristics such as age and gender while preserving speaker identity. In large-scale speech datasets, however, attribute annotations are o…
Audio Language Model for Deepfake Detection Grounded in Acoustic Chain-of-Thought
Runkun Chen, Yixiong Fang, Pengyu Chang +3
Deepfake speech detection systems are often limited to binary classification tasks and struggle to generate interpretable reasoning or provide context-rich explanations for their d…
-DPO: Fairness Direct Preference Optimization Approach to Continual Learning in Large Multimodal Models
Thanh-Dat Truong, Huu-Thien Tran, Jackson Cothren +2
Fairness in Continual Learning for Large Multimodal Models (LMMs) is an emerging yet underexplored challenge, particularly in the presence of imbalanced data distributions that can…
What and When to Learn: CURriculum Ranking Loss for Large-Scale Speaker Verification
Massa Baali, Sarthak Bisht, Rita Singh +1
Speaker verification at large scale remains an open challenge as fixed-margin losses treat all samples equally regardless of quality. We hypothesize that mislabeled or degraded sam…