4 papers
Hierarchical Activity Recognition and Captioning from Long-Form Audio
Peng Zhang, Qingyu Luo, Philip J. B. Jackson +1
Complex activities in real-world audio unfold over extended durations and exhibit hierarchical structure, yet most prior work focuses on short clips and isolated events. To bridge…
Gen-SER: When the generative model meets speech emotion recognition
Taihui Wang, Jinzheng Zhao, Rilin Chen +3
Speech emotion recognition (SER) is crucial in speech understanding and generation. Most approaches are based on either classification models or large language models. Different fr…
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey
Tianxin Xie, Yan Rong, Pengfei Zhang +2
Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over attributes like emotion, timbre, and style. Driven by rising industr…
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems
Huan Zhang, Jinhua Liang, Huy Phan +2
Evaluating generative models remains a fundamental challenge, particularly when the goal is to reflect human preferences. In this paper, we use music generation as a case study to…