1 paper
Kanishk Jain, Qian Yang, Shravan Nayak +3
Vision-language Models (VLMs), despite achieving strong performance on multimodal benchmarks, often misinterpret straightforward visual concepts that humans identify effortlessly,…