4 papers
MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval
Fulong Liu, Liang Xu, Chengqun Yang +3
Human motion-text retrieval provides a rigorous means of assessing cross-modal alignment. Prevailing benchmarks are dominated by homogeneous indoor motions, imbalanced motion distr…
Scaling Up Occupancy-centric Driving Scene Generation: Dataset and Method
Bohan Li, Xin Jin, Hu Zhu +9
Driving scene generation is a critical domain for autonomous driving, enabling downstream applications, including perception and planning evaluation. Occupancy-centric methods have…
OccScene: Semantic Occupancy-based Cross-task Mutual Learning for 3D Scene Generation
Bohan Li, Xin Jin, Jianan Wang +8
Recent diffusion models have demonstrated remarkable performance in both 3D scene generation and perception tasks. Nevertheless, existing methods typically separate these two proce…
Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
Liang Xu, Chengqun Yang, Zili Lin +11
Learning action models from real-world human-centric interaction datasets is important towards building general-purpose intelligent assistants with efficiency. However, most existi…