1 paper
Kazi Hasan Ibn Arif, JinYi Yoon, Dimitrios S. Nikolopoulos +3
High-resolution Vision-Language Models (VLMs) are widely used in multimodal tasks to enhance accuracy by preserving detailed image information. However, these models often generate…