1 paper
Yuanzhi Xu, Qian Gao, Jun Fan +4
Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to the wrong same-class instance. Generic visual-question-answering…