8 papers
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
Yu Zhang, Ruiqi Li, Changhao Pan +3
Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may ne…
EnchantDance: Unveiling the Potential of Music-Driven Dance Movement
Bo Han, Teng Zhang, Zeyu Ling +1
The task of music-driven dance generation involves creating coherent dance movements that correspond to the given music. While existing methods can produce physically plausible dan…
DREAM-R: Multimodal Speculative Reasoning with RL-Based Refined Drafting, Precise Verification, and Fully Parallel Execution
Yunhai Hu, Zining Liu, Xiangyang Yin +5
Speculative reasoning has recently been proposed as a means to accelerate reasoning-intensive generation in large multimodal models, but its effectiveness is often constrained by m…
Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios
Changhao Pan, Rui Yang, Han Wang +12
Recent advances in speech generation have enabled high-fidelity synthesis, yet systematic evaluation of models under long-context conditions remains largely underexplored. A compre…
InfinityHuman: Towards Long-Term Audio-Driven Human
Xiaodi Li, Pan Xie, Yi Ren +6
Audio-driven human animation has attracted wide attention thanks to its practical applications. However, critical challenges remain in generating high-resolution, long-duration vid…
HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation
Qijun Gan, Yi Ren, Chen Zhang +6
Human motion video generation has advanced significantly, while existing methods still struggle with accurately rendering detailed body parts like hands and faces, especially in lo…