6 papers
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
Haitao Lin, Hanyang Yu, Jingshun Huang +5
Existing Vision-Language-Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high-level perception with sparse, embodiment-spec…
MetricHMSR:Metric Human Mesh and Scene Recovery from Monocular Images
Chentao Song, He Zhang, Haolei Yuan +4
We introduce MetricHMSR, a novel framework for recovering metric human meshes and 3D scenes from a single monocular image. Existing methods struggle to recover metric scale due to…
MotionPRO: Exploring the Role of Pressure in Human MoCap and Beyond
Shenghao Ren, Yi Lu, Jiayi Huang +5
Existing human Motion Capture (MoCap) methods mostly focus on the visual similarity while neglecting the physical plausibility. As a result, downstream tasks such as driving virtua…
BioHuman: Learning Biomechanical Human Representations from Video
Yujun Huo, He Zhang, Chentao Song +3
Understanding human motion beyond surface kinematics is crucial for motion analysis, rehabilitation, and injury risk assessment. However, progress in this domain is limited by the…
EgoLive: A Large-Scale Egocentric Dataset from Real-World Human Tasks
Yihang Li, Xuelong Wei, Jingzhou Luo +26
The advancement of robot learning is currently hindered by the scarcity of large-scale, high-quality datasets. While established data collection methods such as teleoperation and u…
DexterCap: An Affordable and Automated System for Capturing Dexterous Hand-Object Manipulation
Yutong Liang, Shiyi Xu, Yulong Zhang +3
Capturing fine-grained hand-object interactions is challenging due to severe self-occlusion from closely spaced fingers and the subtlety of in-hand manipulation motions. Existing o…