1 paper
Yanshu Li, Jianjiang Yang, Zhennan Shen +3
Modern large vision-language models (LVLMs) convert each input image into a large set of tokens that far outnumber the text tokens. Although this improves visual perception, it als…