11 papers
MotionRFT: Unified Reinforcement Fine-Tuning for Text-to-Motion Generation
Xiaofeng Tan, Wanjiang Weng, Hongsong Wang +3
Text-to-motion generation has advanced with diffusion- and flow-based generative models, yet supervised pretraining remains insufficient to align models with high-level objectives…
Temporal Consistency-Aware Text-to-Motion Generation
Hongsong Wang, Wenjing Yan, Qiuxia Lai +1
Text-to-Motion (T2M) generation aims to synthesize realistic human motion sequences from natural language descriptions. While two-stage frameworks leveraging discrete motion repres…
Controllable Dance Generation with Style-Guided Motion Diffusion
Hongsong Wang, Ying Zhu, Xin Geng +1
Dance plays an important role as an artistic form and expression in human culture, yet automatically generating dance sequences is a significant yet challenging endeavor. Existing…
Fast Inference of Visual Autoregressive Model with Adjacency-Adaptive Dynamical Draft Trees
Haodong Lei, Hongsong Wang, Xin Geng +2
Autoregressive (AR) image models achieve diffusion-level quality but suffer from sequential inference, requiring approximately 2,000 steps for a 576x576 image. Speculative decoding…
SoPo: Text-to-Motion Generation Using Semi-Online Preference Optimization
Xiaofeng Tan, Hongsong Wang, Xin Geng +1
Text-to-motion generation is essential for advancing the creative industry but often presents challenges in producing consistent, realistic motions. To address this, we focus on fi…
Foundation Model for Skeleton-Based Human Action Understanding
Hongsong Wang, Wanjiang Weng, Junbo Wang +4
Human action understanding serves as a foundational pillar in the field of intelligent motion perception. Skeletons serve as a modality- and device-agnostic representation for huma…