1 paper
Jin Cui, Yanbin Hu, Xinyue Long +3
Visual representations of VLA models remain unreliable for spatially precise robotic manipulation. We uncover that vision encoders in VLAs also exhibit attention artifacts previous…