activity
20242026
collaborators

5 papers

cs.RO2026

Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation

Yifan Xie, YuAn Wang, Guangyu Chen +3

Human videos contain rich manipulation priors, but using them for robot learning remains difficult because raw observations entangle scene understanding, human motion, and embodime…

cs.CV2025

IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes

Yujia Liang, Jile Jiao, Xuetao Feng +3

Video Large Language Models (VideoLLMs) have demonstrated remarkable understanding capabilities, but are found struggling to tackle multi-shot scenarios,e.g., video clips with vary…

cs.CV2025

Human-Centric Foundation Models: Perception, Generation and Agentic Modeling

Shixiang Tang, Yizhou Wang, Lu Chen +4

Human understanding and generation are critical for modeling digital humans and humanoid embodiments. Recently, Human-centric Foundation Models (HcFMs) inspired by the success of g…

cs.CV2024

MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding

Yuan Wang, Di Huang, Yaqi Zhang +5

Generating lifelike human motions from descriptive texts has experienced remarkable research focus in the recent years, propelled by the emerging requirements of digital humans.Des…

cs.CV2024

Holistic-Motion2D: Scalable Whole-body Human Motion Generation in 2D Space

Yuan Wang, Zhao Wang, Junhao Gong +8

In this paper, we introduce a novel path to human motion generation by focusing on 2D space. Traditional methods have primarily generated human motions in 3D, wh…