6 papers
Spatiotemporally Decoupled Autoregressive Diffusion Model for Human Motion Generation
Chengqun Yang, Liang Xu, Yanping Li +4
Text-driven human motion synthesis has made substantial development with two core modules of motion representation and generative architecture. For representation, Vector Quantizat…
Inter-X++: A Comprehensive Benchmark for Multimodal Human-Human Interaction Analysis
Liang Xu, Chengqun Yang, Zili Lin +6
The capability to perceive and synthesize human-human interactions is fundamental to developing intelligent digital human systems. However, existing datasets and modeling approache…
MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval
Fulong Liu, Liang Xu, Chengqun Yang +3
Human motion-text retrieval provides a rigorous means of assessing cross-modal alignment. Prevailing benchmarks are dominated by homogeneous indoor motions, imbalanced motion distr…
Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control
Weili Zeng, Yitong Xing, Fulong Liu +10
World generative models are typically used through what they produce: a rendered future, a video-conditioned action, or latent context computed by a costly generative branch. We ar…
POLAR: A Portrait OLAT Dataset and Generative Framework for Illumination-Aware Face Modeling
Zhuo Chen, Chengqun Yang, Zhuo Su +5
Face relighting aims to synthesize realistic portraits under novel illumination while preserving identity and geometry. However, progress remains constrained by the limited availab…
Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
Liang Xu, Chengqun Yang, Zili Lin +11
Learning action models from real-world human-centric interaction datasets is important towards building general-purpose intelligent assistants with efficiency. However, most existi…