collaborators

5 papers

cs.SD2026

Fourier is Frontier: Frequency-Aware Autoencoding for High-Fidelity Music Reconstruction

Kangdi Wang, Yusheng Dai, Jin Xu

Continuous-latent audio autoencoders form the backbone of latent music generators, yet decoders at high compression rates commonly exhibit three failure modes: high-frequency loss,…

eess.AS2026

CineDub: Scaling End-to-End Video Dubbing to Multi-Speaker Dialogues with Coherent Sound Effects

Yusheng Dai, Kangdi Wang, Baolong Gao +4

Automatic video dubbing in the wild remains fundamentally limited by two competing constraints: hierarchical methods depend on brittle, multi-stage preprocessing pipelines that sev…

cs.SD2026

Qwen-Music Technical Report

Jin Xu, Kangdi Wang, Ruibin Yuan +24

In this report, we introduce Qwen-Music, a powerful music generation model capable of producing highly musical and high-fidelity songs with complete vocal singing. Qwen-Music suppo…

cs.SD2025

DuoTok: Source-Aware Dual-Track Tokenization for Multi-Track Music Language Modeling

Rui Lin, Zhiyue Wu, Jiahe Le +4

Audio tokenization bridges continuous waveforms and multi-track music language models. In dual-track modeling, tokens should preserve three properties at once: high-fidelity recons…

cs.SD2025

Back to Ear: Perceptually Driven High Fidelity Music Reconstruction

Kangdi Wang, Zhiyue Wu, Dinghao Zhou +3

Variational Autoencoders (VAEs) are essential for large-scale audio tasks like diffusion-based generation. However, existing open-source models often neglect auditory perceptual as…