3 papers
cs.RO2026
Direct Action-Head Injection of A Grounded 3D Point Unlocks Spatial and Task Generalization
Shiang-Feng Tsai, Jin-Cheng Jhang, Yen-Ling Tai +4
Vision-Language-Action (VLA) models leverage large-scale vision-language pretraining for flexible robot manipulation, yet at test time they remain brittle along two axes: spatial g…
cs.RO2026
GRITS: A Spillage-Aware Guided Diffusion Policy for Robot Food Scooping Tasks
Yen-Ling Tai, Yi-Ru Yang, Kuan-Ting Yu +2
Robotic food scooping is a critical manipulation skill for food preparation and service robots. However, existing robot learning algorithms, especially learn-from-demonstration met…
cs.RO2025
Mitigating Cross-Modal Distraction and Ensuring Geometric Feasibility via Affordance-Guided and Self-Consistent MLLMs for Task Planning in Instruction-Following Manipulation
Yu-Hong Shen, Chuan-Yu Wu, Yi-Ru Yang +2
We investigate the use of Multimodal Large Language Models (MLLMs) with in-context learning for closed-loop task planning in instruction-following manipulation. We identify four es…