5 papers
CLASVS: Continuous-Latent Autoregression for Melody-Preserving Lyric Editing in Singing Voice Synthesis
Yizhong Geng, Tian-Hao Zhang, Chunfeng Wang +7
Reference-conditioned melody-preserving lyric editing replaces words while retaining a performance's timing, singer identity, and naturalness. Continuous-latent autoregression avoi…
ChainSpace: A Chained-Reasoning Paradigm for Spatial Intelligence
Xiaohan Zhang, Feng Gu, Xudong Rao +4
Spatial intelligence requires foundation models to maintain coherent spatial state across interactions with the physical world. However, existing data-centric approaches typically…
Towards More Expressive Spoken LLMs: Fine-Grained Intent Benchmarking and Acoustic-Lexical Decoupled Policy Optimization
Xiang Lin, Tian-Hao Zhang, Chunfeng Wang +3
Spoken emotional dialogue requires a model to understand a user's spoken input and generate a response that is both semantically appropriate and emotionally expressive. This is cha…
PALM-Bench: A Comprehensive Benchmark for Personalized Audio-Language Models
Yuwen Wang, Xinyuan Qian, Tian-Hao Zhang +6
Large Audio-Language Models (LALMs) have demonstrated strong performance in audio understanding and generation. Yet, our extensive benchmarking reveals that their behavior is large…
MSU-Bench: Towards Understanding the Conversational Multi-talker Scenarios
Shuai Wang, Zhaokai Sun, Zhennan Lin +3
Spoken Language Understanding (SLU) has progressed from traditional single-task methods to large audio language model (LALM) solutions. Yet, most existing speech benchmarks focus o…