1 paper
Chenrui Fan, Yijun Liang, Shweta Bhardwaj +3
While many vision-language models (VLMs) are developed to answer well-defined, straightforward questions with highly specified targets, as in most benchmarks, they often struggle i…