1 paper · 1 filter
Jinsik Bang, Jaeyeon Bae, Donggyu Lee +2
Vision-language models (VLMs) have shown strong perception and reasoning abilities for instruction-following embodied agents. However, despite these abilities and their generalizat…