8 papers
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens
Yi Luo, Rongzhi Gu, Jixun Yao
Balancing sequence length, representational capacity, and long-horizon stability is a central problem in autoregressive (AR) speech and audio generation. Representations with highe…
SongPrep: A Preprocessing Framework and End-to-end Model for Full-song Structure Parsing and Lyrics Transcription
Wei Tan, Shun Lei, Huaicheng Zhang +6
Artificial Intelligence Generated Content (AIGC) is currently a popular research area. Among its various branches, song generation has attracted growing interest. Despite the abund…
MuCodec: Ultra Low-Bitrate Music Codec
Yaoxun Xu, Hangting Chen, Jianwei Yu +5
Music codecs are a vital aspect of audio codec research, and ultra low-bitrate compression holds significant importance for music transmission and generation. Due to the complexity…
WAKE: Watermarking Audio with Key Enrichment
Yaoxun Xu, Jianwei Yu, Hangting Chen +5
As deep learning advances in audio generation, challenges in audio security and copyright protection highlight the need for robust audio watermarking. Recent neural network-based m…
SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor
Chenyu Yang, Shuai Wang, Hangting Chen +7
The emergence of novel generative modeling paradigms, particularly audio language models, has significantly advanced the field of song generation. Although state-of-the-art models…
MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
Haina Zhu, Yizhi Zhou, Hangting Chen +6
Recent years have witnessed the success of foundation models pre-trained with self-supervised learning (SSL) in various music informatics understanding tasks, including music taggi…