1 paper
Sangoh Lee, Sangwoo Mo, Wook-Shin Han
While Vision-Language-Action (VLA) models generalize well to generic instructions, they struggle with personalized commands such as "bring my cup," where the robot must act on one…