1 paper · 1 filter
Xiaowen Sun, Matthias Kerzel, Mengdi Li +3
Vision-language models (VLMs) have shown remarkable performance in various robotic tasks, as they can perceive visual information and understand natural language instructions. Howe…