1 paper
Xiaowen Sun, Matthias Kerzel, Mengdi Li +3
Vision-language models (VLMs) have shown remarkable performance in various robotic tasks, as they can perceive visual information and understand natural language instructions. Howe…