6 papers
GWM-VLA: Geometry-Aware Latent World Modeling for Vision-Language-Action Learning
Yanping Zhao, Hang Yu, Yiwei Wang +7
Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but often degrade under visual and environmental shifts. Latent world modeling offers a promisin…
ACSAC: Adaptive Chunk Size Actor-Critic with Causal Transformer Q-Network
Qian Chen, Junqiao Zhao, Hongtu Zhou +4
Long-horizon, sparse-reward tasks pose a fundamental challenge for reinforcement learning, since single-step TD learning suffers from bootstrapping error accumulation across succes…
Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning
Qingjun Wang, Hongtu Zhou, Hang Yu +5
Offline reinforcement learning (RL) faces a critical challenge of overestimating the value of out-of-distribution (OOD) actions. Existing methods mitigate this issue by penalizing…
Point What You Mean: Visually Grounded Instruction Policy
Hang Yu, Juntu Zhao, Yufeng Liu +9
Vision-Language-Action (VLA) models align vision and language with embodied control, but their object referring ability remains limited when relying solely on text prompt, especial…
Learning Geometrically-Grounded 3D Visual Representations for View-Generalizable Robotic Manipulation
Di Zhang, Weicheng Duan, Dasen Gu +5
Real-world robotic manipulation demands visuomotor policies capable of robust spatial scene understanding and strong generalization across diverse camera viewpoints. While recent a…
ASTRO: Adaptive Stitching via Dynamics-Guided Trajectory Rollouts
Hang Yu, Di Zhang, Qiwei Du +5
Offline reinforcement learning (RL) enables agents to learn optimal policies from pre-collected datasets. However, datasets containing suboptimal and fragmented trajectories presen…