1 paper
Wenbo Hu, Zi-Yi Dou, Liunian Harold Li +3
Large Vision-Language Models (LVLMs) typically encode an image into a fixed number of visual tokens (e.g., 576) and process these tokens with a language model. Despite their strong…