1 paper · 1 filter
Rory Driscoll, Alexandros Christoforos, Chadbourne Davis
While sequential reasoning enhances the capability of Vision-Language Models (VLMs) to execute complex multimodal tasks, their reliability in grounding these reasoning chains withi…