13 papers
Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering
Junyu Dai, Xinyue Fan, Weiqin Li +14
In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descriptions, and musical attributes. The propo…
Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models
Yizhou Peng, Yukun Ma, Chong Zhang +4
While Text-to-Speech (TTS) systems enable emotional control via natural-language instructions, expressiveness, naturalness, and speech quality degrade when the target emotion confl…
LuSeeL: Language-queried Binaural Universal Sound Event Extraction and Localization
Zexu Pan, Shengkui Zhao, Yukun Ma +4
Most universal sound extraction algorithms focus on isolating a target sound event from single-channel audio mixtures. However, the real world is three-dimensional, and binaural au…
FGGM: Fisher-Guided Gradient Masking for Continual Learning
Chao-Hong Tan, Qian Chen, Wen Wang +6
Catastrophic forgetting impairs the continuous learning of large language models. We propose Fisher-Guided Gradient Masking (FGGM), a framework that mitigates this by strategically…
Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions
Kun Zhou, You Zhang, Dianwen Ng +3
Emotional text-to-speech (TTS) systems sturggle to capture the full spectrum of human emotions due to the inherent complexity of emotional expressions and the limited coverage of e…
DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations
Chao-Hong Tan, Qian Chen, Wen Wang +14
Recent studies on end-to-end (E2E) speech generation with large language models (LLMs) have attracted significant community attention, with multiple works extending text-based LLMs…