3 citations · 8 across the 23 of their papers we have counts for
1 paper · 2 filters
Di Zhang, Junxian Li, Jingdi Lei +10
Vision-language models (VLMs) have shown remarkable advancements in multimodal reasoning tasks. However, they still often generate inaccurate or irrelevant responses due to issues…