1 paper
Hyunsik Chae, Seungwoo Yoon, Jaden Park +5
Recent Vision-Language Models (VLMs) have demonstrated impressive multimodal comprehension and reasoning capabilities, yet they often struggle with trivially simple visual tasks. I…