2 papers
cs.RO2026
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
Zhengshen Zhang, Hao Li, Yalun Dai +10
Existing vision-language-action (VLA) models act in 3D real-world but are typically built on 2D encoders, leaving a spatial reasoning gap that limits generalization and adaptabilit…
cs.RO2025
3D Affordance Keypoint Detection for Robotic Manipulation
Zhiyang Liu, Ruiteng Zhao, Lei Zhou +6
This paper presents a novel approach for affordance-informed robotic manipulation by introducing 3D keypoints to enhance the understanding of object parts' functionality. The propo…