4 papers
Hierarchical Activity Recognition and Captioning from Long-Form Audio
Peng Zhang, Qingyu Luo, Philip J. B. Jackson +1
Complex activities in real-world audio unfold over extended durations and exhibit hierarchical structure, yet most prior work focuses on short clips and isolated events. To bridge…
Gen-SER: When the generative model meets speech emotion recognition
Taihui Wang, Jinzheng Zhao, Rilin Chen +3
Speech emotion recognition (SER) is crucial in speech understanding and generation. Most approaches are based on either classification models or large language models. Different fr…
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems
Huan Zhang, Jinhua Liang, Huy Phan +2
Evaluating generative models remains a fundamental challenge, particularly when the goal is to reflect human preferences. In this paper, we use music generation as a case study to…
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey
Tianxin Xie, Yan Rong, Pengfei Zhang +2
Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over attributes like emotion, timbre, and style. Driven by rising industr…