collaborators

5 papers

cs.LG2026

Encoder-Decoder Manifold Alignment for Idempotent Generation

Dareen Alharthi, Abdul Waheed, Bhiksha Raj

Recently, several learning paradigms have been introduced to enforce idempotency in generative models. The goal is to ensure that repeated application of a model leaves samples unc…

cs.SD2026

RIVET: Robust Idempotent Voice Attribute Editing

Dareen Alharthi, Bhuvan Koduru, Rita Singh +1

Voice attribute editing models modify characteristics such as age and gender while preserving speaker identity. In large-scale speech datasets, however, attribute annotations are o…

cs.CV2025

VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding

Abdul Waheed, Zhen Wu, Dareen Alharthi +2

Precisely evaluating video understanding models remains challenging: commonly used metrics such as BLEU, ROUGE, and BERTScore fail to capture the fineness of human judgment, while…

cs.SD2025

VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music

Jiatong Shi, Hye-jin Shim, Jinchuan Tian +14

In this work, we introduce VERSA, a unified and standardized evaluation toolkit designed for various speech, audio, and music signals. The toolkit features a Pythonic interface wit…

cs.LG2025

Tessellated Linear Model for Age Prediction from Voice

Dareen Alharthi, Mahsa Zamani, Bhiksha Raj +1

Voice biometric tasks, such as age estimation require modeling the often complex relationship between voice features and the biometric variable. While deep learning models can hand…