1 paper
Nishanth Kumar, William Shen, Fabio Ramos +4
Foundation models like Vision-Language Models (VLMs) excel at common sense vision and language tasks such as visual question answering. However, they cannot yet directly solve comp…