5 papers
AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling
Yiming Ren, Xuenan Xu, Ziyang Zhang +3
Despite advances in text and visual generation, creating coherent long-form audio narratives remains challenging. Existing frameworks often exhibit limitations such as mismatched c…
HoliAntiSpoof: Audio LLM for Holistic Speech Anti-Spoofing
Xuenan Xu, Yiming Ren, Liwei Liu +5
Recent advances in speech synthesis and editing have made speech spoofing increasingly challenging. However, most existing methods treat spoofing as binary classification, overlook…
Photo- Estimation with Normalizing Flow
Yiming Ren, Kwan Chuen Chan, Le Zhang +6
Accurate photometric redshift (photo-) estimation is a key challenge in cosmology, as uncertainties in photo- directly limit the scientific return of large-scale structure an…
Can Audio Large Language Models Verify Speaker Identity?
Yiming Ren, Xuenan Xu, Baoxiang Li +2
This paper investigates adapting Audio Large Language Models (ALLMs) for speaker verification (SV). We reformulate SV as an audio question-answering task and conduct comprehensive…
Debiasing Multimodal Large Language Models via Noise-Aware Preference Optimization
Zefeng Zhang, Hengzhu Tang, Jiawei Sheng +6
Multimodal Large Language Models excel in various tasks, yet often struggle with modality bias, where the model tends to rely heavily on a single modality and overlook critical inf…