1 paper · 1 filter
Junjie Chen, Xuyang Liu, Zichen Wen +3
Large vision-language models (LVLMs) have demonstrated remarkable capabilities in multimodal understanding tasks. However, the increasing demand for high-resolution image and long-…