Showing cs.ROShow all
3 papers · 1 filter
cs.RO2026
KAM-WM: Kinematic Affordance Maps from Latent World Models for Robot Manipulation
Xinyu Shao, Keru Zhou, Guowei Huang +3
Learning manipulation from few demonstrations requires visual priors that capture not only where to interact, but also how the interaction should begin; static priors such as segme…
cs.RO2026
Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning
Yanzhe Tang, Xinyu Shao, Yuxuan Hu +6
While end-to-end Vision-Language-Action (VLA) models show promise in robotic manipulation, their monolithic paradigm inherently couples semantic reasoning and spatial control. This…
cs.RO2026
ELAN4D: Embodiment-Centric 4D Supervision for Vision-Language-Action Models via Plug-and-Play Adaptation
Zeyuan He, Bowen Yang, Zhirui Fang +9
Vision-Language-Action (VLA) models have shown promise for robotic manipulation, yet most existing policies operate reactively by directly regressing actions from current observati…