4 citations · 27 across the 23 of their papers we have counts for
1 paper · 2 filters
Zipeng Zhu, Zhanghao Hu, Qinglin Zhu +5
Large Vision-Language Models (LVLMs) have advanced rapidly by aligning visual patches with the text embedding space, but a fixed visual-token budget forces images to be resized to…