3 papers
cs.RO2026
A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model
Kaidong Zhang, Jian Zhang, Rongtao Xu +20
Vision-Language-Action (VLA) models have emerged as a powerful paradigm for open-world robot manipulation, but their practical deployment is often constrained by cost: billion-scal…
cs.CV2026
Egocentric Visibility-Aware Human Pose Estimation
Peng Dai, Yu Zhang, Yiqiang Feng +2
Egocentric human pose estimation (HPE) using a head-mounted device is crucial for various VR and AR applications, but it faces significant challenges due to keypoint invisibility.…
cs.CV2025
InterPose: Learning to Generate Human-Object Interactions from Large-Scale Web Videos
Yangsong Zhang, Abdul Ahad Butt, Gül Varol +1
Human motion generation has shown great advances thanks to the recent diffusion models trained on large-scale motion capture data. Most of existing works, however, currently target…