12 papers
GWM-VLA: Geometry-Aware Latent World Modeling for Vision-Language-Action Learning
Yanping Zhao, Hang Yu, Yiwei Wang +7
Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but often degrade under visual and environmental shifts. Latent world modeling offers a promisin…
SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
Peizheng Li, Zhenghao Zhang, David Holtz +6
End-to-end autonomous driving methods built on vision language models (VLMs) have undergone rapid development driven by their universal visual understanding and strong reasoning ca…
TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging
Shengzhuo Yang, Ronghao Yu, Chuanjie Lv +5
Vision-language-action (VLA) models aim to understand natural-language instructions and visual observations, and to generate and execute corresponding actions as embodied agents. R…
Learning Native Continuation for Action Chunking Flow Policies
Yufeng Liu, Hang Yu, Juntu Zhao +9
Action chunking enables Vision Language Action (VLA) models to run in real time, but naive chunked execution often exhibits discontinuities at chunk boundaries. Real-Time Chunking…
Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning
Qingjun Wang, Hongtu Zhou, Hang Yu +5
Offline reinforcement learning (RL) faces a critical challenge of overestimating the value of out-of-distribution (OOD) actions. Existing methods mitigate this issue by penalizing…
Point What You Mean: Visually Grounded Instruction Policy
Hang Yu, Juntu Zhao, Yufeng Liu +9
Vision-Language-Action (VLA) models align vision and language with embodied control, but their object referring ability remains limited when relying solely on text prompt, especial…