8 papers
CLASVS: Continuous-Latent Autoregression for Melody-Preserving Lyric Editing in Singing Voice Synthesis
Yizhong Geng, Tian-Hao Zhang, Chunfeng Wang +7
Reference-conditioned melody-preserving lyric editing replaces words while retaining a performance's timing, singer identity, and naturalness. Continuous-latent autoregression avoi…
Towards More Expressive Spoken LLMs: Fine-Grained Intent Benchmarking and Acoustic-Lexical Decoupled Policy Optimization
Xiang Lin, Tian-Hao Zhang, Chunfeng Wang +3
Spoken emotional dialogue requires a model to understand a user's spoken input and generate a response that is both semantically appropriate and emotionally expressive. This is cha…
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models
Yuwen Wang, Tian-Hao Zhang, Minghao Cai +7
Complex acoustic problems may require models to perform acoustic operations, interact with external tools and reason over the resulting textual or processed-audio observations rath…
PALM-Bench: A Comprehensive Benchmark for Personalized Audio-Language Models
Yuwen Wang, Xinyuan Qian, Tian-Hao Zhang +6
Large Audio-Language Models (LALMs) have demonstrated strong performance in audio understanding and generation. Yet, our extensive benchmarking reveals that their behavior is large…
IKFST: IOO and KOO Algorithms for Accelerated and Precise WFST-based End-to-End Automatic Speech Recognition
Zhuoran Zhuang, Ye Chen, Chao Luo +5
End-to-end automatic speech recognition has become the dominant paradigm in both academia and industry. To enhance recognition performance, the Weighted Finite-State Transducer (WF…
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception
Jiawei Zhang, Tian-Hao Zhang, Jun Wang +3
Controlling the style and characteristics of speech synthesis is crucial for adapting the output to specific contexts and user requirements. Previous Text-to-speech (TTS) works hav…