activity
20242026
collaborators

35 papers

cs.LG2026

Encoder-Decoder Manifold Alignment for Idempotent Generation

Dareen Alharthi, Abdul Waheed, Bhiksha Raj

Recently, several learning paradigms have been introduced to enforce idempotency in generative models. The goal is to ensure that repeated application of a model leaves samples unc…

cs.SD2026

RIVET: Robust Idempotent Voice Attribute Editing

Dareen Alharthi, Bhuvan Koduru, Rita Singh +1

Voice attribute editing models modify characteristics such as age and gender while preserving speaker identity. In large-scale speech datasets, however, attribute annotations are o…

cs.SD2026

Audio Language Model for Deepfake Detection Grounded in Acoustic Chain-of-Thought

Runkun Chen, Yixiong Fang, Pengyu Chang +3

Deepfake speech detection systems are often limited to binary classification tasks and struggle to generate interpretable reasoning or provide context-rich explanations for their d…

cs.LG2026

-DPO: Fairness Direct Preference Optimization Approach to Continual Learning in Large Multimodal Models

Thanh-Dat Truong, Huu-Thien Tran, Jackson Cothren +2

Fairness in Continual Learning for Large Multimodal Models (LMMs) is an emerging yet underexplored challenge, particularly in the presence of imbalanced data distributions that can…

cs.SD2026

What and When to Learn: CURriculum Ranking Loss for Large-Scale Speaker Verification

Massa Baali, Sarthak Bisht, Rita Singh +1

Speaker verification at large scale remains an open challenge as fixed-margin losses treat all samples equally regardless of quality. We hypothesize that mislabeled or degraded sam…

cs.SD2026

DELULU: Discriminative Embedding Learning Using Latent Units for Speaker-Aware Self-Trained Speech Foundational Model

Massa Baali, Rita Singh, Bhiksha Raj

Self-supervised speech models have achieved remarkable success on content-driven tasks, yet they remain limited in capturing speaker-discriminative features critical for verificati…