1 paper
Jiawei Guo, Feifei Zhai, Pu Jian +2
Current VLM-based VQA methods often process entire images, leading to excessive visual tokens that include redundant information irrelevant to the posed question. This abundance of…