1 paper · 1 filter
Chengyi Du, Yazhe Niu, Dazhong Shen +1
Recent advances in vision-language models (VLMs) have markedly improved image-text alignment, yet they still fall short of human-like visual reasoning. A key limitation is that man…