1 paper · 1 filter
Ashish Baghel, Paras Chopra
Vision-Language Models (VLMs) excel at describing visual scenes, yet struggle to translate perception into precise, grounded actions. We investigate whether providing VLMs with bot…