1 paper
Guanhua Ye, Niu Jingbin, Yan Li +4
Short-answer VQA benchmarks conflate two distinct quantities: whether a model's answer is semantically correct, and whether that answer matches the surface form expected by the aut…