1 paper
Feixiang Liu, Qiang Qiu, Qingyang Li +1
Vision-language models can answer spatial relation questions confidently even when the image supports an incompatible relation. We formulate relation-grounded selective prediction:…