6 citations · 6 across the 3 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
CROP: Contextual Region-Oriented Visual Token Pruning
Jiawei Guo, Feifei Zhai, Pu Jian +2
Current VLM-based VQA methods often process entire images, leading to excessive visual tokens that include redundant information irrelevant to the posed question. This abundance of…
cs.CV2025
Teaching Vision-Language Models to Ask: Resolving Ambiguity in Visual Questions
Pu Jian, Donglei Yu, Wen Yang +2
In visual question answering (VQA) context, users often pose ambiguous questions to visual language models (VLMs) due to varying expression habits. Existing research addresses such…
cs.CV2025
Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
Pu Jian, Junhong Wu, Wei Sun +3
Recent advances in text-only "slow-thinking" reasoning have prompted efforts to transfer this capability to vision-language models (VLMs), for training visual reasoning models (\te…