3 papers
cs.SD2026
TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation
Xiaoda Yang, Majun Zhang, Changhao Pan +10
Unified audio-visual generation is rapidly gaining industrial and creative relevance, enabling applications in virtual production and interactive media. However, when moving from g…
cs.CV2026
Not All Frames Are Equal: Complexity-Aware Masked Motion Generation via Motion Spectral Descriptors
Pengfei Zhou, Xiangyue Zhang, Xukun Shen +1
Masked generative models have become a strong paradigm for text-to-motion synthesis, but they still treat motion frames too uniformly during masking, attention, and decoding. This…
cs.CV2025
Text-driven 3D Human Generation via Contrastive Preference Optimization
Pengfei Zhou, Xukun Shen, Yong Hu
Recent advances in Score Distillation Sampling (SDS) have improved 3D human generation from textual descriptions. However, existing methods still face challenges in accurately alig…