3 papers
cs.SD2026
VIBEVOICE-ASR Technical Report
Zhiliang Peng, Jianwei Yu, Yaoyao Chang +21
This report presents VibeVoice-ASR, a general-purpose speech understanding framework built upon VibeVoice, designed to address the persistent challenges of context fragmentation an…
cs.SD2026
The SJTU X-LANCE Lab System for MSR Challenge 2025
Jinxuan Zhu, Hao Qiu, Haina Zhu +3
This report describes the system submitted to the music source restoration (MSR) Challenge 2025. Our approach is composed of sequential BS-RoFormers, each dealing with a single tas…
cs.SD2025
MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
Haina Zhu, Yizhi Zhou, Hangting Chen +6
Recent years have witnessed the success of foundation models pre-trained with self-supervised learning (SSL) in various music informatics understanding tasks, including music taggi…