1 paper · 1 filter
Michael Ogezi, Freda Shi
Vision-language models (VLMs) work well in tasks ranging from image captioning to visual question answering (VQA), yet they struggle with spatial reasoning, a key skill for underst…