1 paper
Gengyuan Zhang, Yurui Zhang, Kerui Zhang +1
Vision-Language Models (VLMs) are expected to be capable of reasoning with commonsense knowledge as human beings. One example is that humans can reason where and when an image is t…