6 papers · 1 filter
GWM-VLA: Geometry-Aware Latent World Modeling for Vision-Language-Action Learning
Yanping Zhao, Hang Yu, Yiwei Wang +7
Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but often degrade under visual and environmental shifts. Latent world modeling offers a promisin…
TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging
Shengzhuo Yang, Ronghao Yu, Chuanjie Lv +5
Vision-language-action (VLA) models aim to understand natural-language instructions and visual observations, and to generate and execute corresponding actions as embodied agents. R…
Learning Native Continuation for Action Chunking Flow Policies
Yufeng Liu, Hang Yu, Juntu Zhao +9
Action chunking enables Vision Language Action (VLA) models to run in real time, but naive chunked execution often exhibits discontinuities at chunk boundaries. Real-Time Chunking…
Robust and Generalized Humanoid Motion Tracking
Yubiao Ma, Han Yu, Jiayin Xie +7
Learning a general humanoid whole-body controller is challenging because practical reference motions can exhibit noise and inconsistencies after being transferred to the robot doma…
Learning Geometrically-Grounded 3D Visual Representations for View-Generalizable Robotic Manipulation
Di Zhang, Weicheng Duan, Dasen Gu +5
Real-world robotic manipulation demands visuomotor policies capable of robust spatial scene understanding and strong generalization across diverse camera viewpoints. While recent a…
DexH2R: A Benchmark for Dynamic Dexterous Grasping in Human-to-Robot Handover
Youzhuo Wang, Jiayi Ye, Chuyang Xiao +6
Handover between a human and a dexterous robotic hand is a fundamental yet challenging task in human-robot collaboration. It requires handling dynamic environments and a wide varie…