46 citations · 66 across the 11 of their papers we have counts for
1 paper · 1 filter
Di Zhang, Junxian Li, Jingdi Lei +10
Vision-language models (VLMs) have shown remarkable advancements in multimodal reasoning tasks. However, they still often generate inaccurate or irrelevant responses due to issues…