14 papers
MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval
Fulong Liu, Liang Xu, Chengqun Yang +3
Human motion-text retrieval provides a rigorous means of assessing cross-modal alignment. Prevailing benchmarks are dominated by homogeneous indoor motions, imbalanced motion distr…
Fine-Grained Human Pose Editing Assessment via Layer-Selective MLLMs
Ningyu Sun, Zhaolin Cai, Zitong Xu +5
Text-guided human pose editing has gained significant traction in AIGC applications. However,it remains plagued by structural anomalies and generative artifacts. Existing evaluatio…
SingingBot: An Avatar-Driven System for Robotic Face Singing Performance
Zhuoxiong Xu, Xuanchen Li, Yuhao Cheng +3
Equipping robotic faces with singing capabilities is crucial for empathetic Human-Robot Interaction. However, existing robotic face driving research primarily focuses on conversati…
POLAR: A Portrait OLAT Dataset and Generative Framework for Illumination-Aware Face Modeling
Zhuo Chen, Chengqun Yang, Zhuo Su +5
Face relighting aims to synthesize realistic portraits under novel illumination while preserving identity and geometry. However, progress remains constrained by the limited availab…
MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts
Jingnan Gao, Zhe Wang, Xianze Fang +7
Recent advances in language and vision have demonstrated that scaling up model capacity consistently improves performance across diverse tasks. In 3D visual geometry reconstruction…
Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
Liang Xu, Chengqun Yang, Zili Lin +11
Learning action models from real-world human-centric interaction datasets is important towards building general-purpose intelligent assistants with efficiency. However, most existi…