1 citations · 1 across the 31 of their papers we have counts for
1 paper · 1 filter
Haiwen Li, Jing Tang, Rui Chen +2
Vision Language Models (VLMs) have demonstrated remarkable capabilities in multimodal reasoning tasks, yet they still suffer from recurring failures, such as skipping key visual ch…