1 paper · 1 filter
Zicong Tang, Ziyang Ma, Suqing Wang +5
Large Vision-Language Models (LVLMs) process multimodal inputs consisting of text tokens and vision tokens extracted from images or videos. Due to the rich visual information, a si…