1 paper · 1 filter
Sangyun Park, Jin Kim, Yuchen Cui +1
Vision-Language Models (VLMs) struggle to translate high-level instructions into the precise spatial affordances required for robotic manipulation. While visual Chain-of-Thought (C…