4 papers · 1 filter
See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models
Byungkun Lee, Dongyoon Hwang, Dongjin Kim +3
Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in the robot's own 3D coordinate frame, yet…
3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance
Dongyoon Hwang, Byungkun Lee, Dongjin Kim +7
Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve generalization in robot manipulation. Recent work in this paradigm u…
Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement
Kinam Kim, Namiko Saito, Heecheol Kim +3
Vision-Language-Action (VLA) models can generalize across diverse manipulation tasks, but their imitation-learning-based policies remain brittle in precise physical interactions du…
ACG: Action Coherence Guidance for Flow-based Vision-Language-Action models
Minho Park, Kinam Kim, Junha Hyung +5
Diffusion and flow matching models have emerged as powerful robot policies, enabling Vision-Language-Action (VLA) models to generalize across diverse scenes and instructions. Yet,…