5 papers
GWM-VLA: Geometry-Aware Latent World Modeling for Vision-Language-Action Learning
Yanping Zhao, Hang Yu, Yiwei Wang +7
Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but often degrade under visual and environmental shifts. Latent world modeling offers a promisin…
Learning Geometrically-Grounded 3D Visual Representations for View-Generalizable Robotic Manipulation
Di Zhang, Weicheng Duan, Dasen Gu +5
Real-world robotic manipulation demands visuomotor policies capable of robust spatial scene understanding and strong generalization across diverse camera viewpoints. While recent a…
ASTRO: Adaptive Stitching via Dynamics-Guided Trajectory Rollouts
Hang Yu, Di Zhang, Qiwei Du +5
Offline reinforcement learning (RL) enables agents to learn optimal policies from pre-collected datasets. However, datasets containing suboptimal and fragmented trajectories presen…
KineDex: Learning Tactile-Informed Visuomotor Policies via Kinesthetic Teaching for Dexterous Manipulation
Di Zhang, Chengbo Yuan, Chuan Wen +3
Collecting demonstrations enriched with fine-grained tactile information is critical for dexterous manipulation, particularly in contact-rich tasks that require precise force contr…
Focus On What Matters: Separated Models For Visual-Based RL Generalization
Di Zhang, Bowen Lv, Hai Zhang +7
A primary challenge for visual-based Reinforcement Learning (RL) is to generalize effectively across unseen environments. Although previous studies have explored different auxiliar…