3 papers
cs.CL2026
PALM-Bench: A Comprehensive Benchmark for Personalized Audio-Language Models
Yuwen Wang, Xinyuan Qian, Tian-Hao Zhang +6
Large Audio-Language Models (LALMs) have demonstrated strong performance in audio understanding and generation. Yet, our extensive benchmarking reveals that their behavior is large…
cs.SD2025
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception
Jiawei Zhang, Tian-Hao Zhang, Jun Wang +3
Controlling the style and characteristics of speech synthesis is crucial for adapting the output to specific contexts and user requirements. Previous Text-to-speech (TTS) works hav…
cs.SD2025
SAV-SE: Scene-aware Audio-Visual Speech Enhancement with Selective State Space Model
Xinyuan Qian, Jiaran Gao, Yaodan Zhang +4
Speech enhancement plays an essential role in various applications, and the integration of visual information has been demonstrated to bring substantial advantages. However, the ma…