1 paper
Yuna Lee, Kyoungho Min, Yulhwa Kim
Recent advancements in Vision-Language Models (VLMs) enable large language models (LLMs) to process high-resolution images, significantly improving real-world multimodal understand…