5 papers
Fourier is Frontier: Frequency-Aware Autoencoding for High-Fidelity Music Reconstruction
Kangdi Wang, Yusheng Dai, Jin Xu
Continuous-latent audio autoencoders form the backbone of latent music generators, yet decoders at high compression rates commonly exhibit three failure modes: high-frequency loss,…
CineDub: Scaling End-to-End Video Dubbing to Multi-Speaker Dialogues with Coherent Sound Effects
Yusheng Dai, Kangdi Wang, Baolong Gao +4
Automatic video dubbing in the wild remains fundamentally limited by two competing constraints: hierarchical methods depend on brittle, multi-stage preprocessing pipelines that sev…
Qwen-Music Technical Report
Jin Xu, Kangdi Wang, Ruibin Yuan +24
In this report, we introduce Qwen-Music, a powerful music generation model capable of producing highly musical and high-fidelity songs with complete vocal singing. Qwen-Music suppo…
DuoTok: Source-Aware Dual-Track Tokenization for Multi-Track Music Language Modeling
Rui Lin, Zhiyue Wu, Jiahe Le +4
Audio tokenization bridges continuous waveforms and multi-track music language models. In dual-track modeling, tokens should preserve three properties at once: high-fidelity recons…
Back to Ear: Perceptually Driven High Fidelity Music Reconstruction
Kangdi Wang, Zhiyue Wu, Dinghao Zhou +3
Variational Autoencoders (VAEs) are essential for large-scale audio tasks like diffusion-based generation. However, existing open-source models often neglect auditory perceptual as…