1 paper
Yiyang Chen, Yixin Tan, Binrui Shen
Recent work typically assesses vision--language consistency using attention distributions of answer-side tokens. However, we observe that highest attention regions are not always c…