12 citations · 12 across the 1 of their papers we have counts for
1 paper · 1 filter
Jiawei Guo, Feifei Zhai, Pu Jian +2
Current VLM-based VQA methods often process entire images, leading to excessive visual tokens that include redundant information irrelevant to the posed question. This abundance of…