35 papers
Encoder-Decoder Manifold Alignment for Idempotent Generation
Dareen Alharthi, Abdul Waheed, Bhiksha Raj
Recently, several learning paradigms have been introduced to enforce idempotency in generative models. The goal is to ensure that repeated application of a model leaves samples unc…
RIVET: Robust Idempotent Voice Attribute Editing
Dareen Alharthi, Bhuvan Koduru, Rita Singh +1
Voice attribute editing models modify characteristics such as age and gender while preserving speaker identity. In large-scale speech datasets, however, attribute annotations are o…
Audio Language Model for Deepfake Detection Grounded in Acoustic Chain-of-Thought
Runkun Chen, Yixiong Fang, Pengyu Chang +3
Deepfake speech detection systems are often limited to binary classification tasks and struggle to generate interpretable reasoning or provide context-rich explanations for their d…
-DPO: Fairness Direct Preference Optimization Approach to Continual Learning in Large Multimodal Models
Thanh-Dat Truong, Huu-Thien Tran, Jackson Cothren +2
Fairness in Continual Learning for Large Multimodal Models (LMMs) is an emerging yet underexplored challenge, particularly in the presence of imbalanced data distributions that can…
What and When to Learn: CURriculum Ranking Loss for Large-Scale Speaker Verification
Massa Baali, Sarthak Bisht, Rita Singh +1
Speaker verification at large scale remains an open challenge as fixed-margin losses treat all samples equally regardless of quality. We hypothesize that mislabeled or degraded sam…
DELULU: Discriminative Embedding Learning Using Latent Units for Speaker-Aware Self-Trained Speech Foundational Model
Massa Baali, Rita Singh, Bhiksha Raj
Self-supervised speech models have achieved remarkable success on content-driven tasks, yet they remain limited in capturing speaker-discriminative features critical for verificati…