Showing cs.ROShow all
2 papers · 1 filter
cs.RO2026
Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models
Sangoh Lee, Sangwoo Mo, Wook-Shin Han
Vision-Language-Action (VLA) models can turn multimodal context into robot actions, but their action decoders are still trained largely by behavior cloning. This supervises which m…
cs.RO2026
Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting
Sangoh Lee, Sangwoo Mo, Wook-Shin Han
While Vision-Language-Action (VLA) models generalize well to generic instructions, they struggle with personalized commands such as "bring my cup," where the robot must act on one…