1 paper
Shulin Tian, Minglun Li, Yuhao Dong +6
Vision-language models can produce fluent answers that are insufficiently grounded in the visual evidence: a single unsupported object, chart value, or intermediate inference can u…