9 citations · 9 across the 2 of their papers we have counts for
2 papers
cs.CV2025
Attention Debiasing for Token Pruning in Vision Language Models
Kai Zhao, Wubang Yuan, Yuchen Lin +5
Vision-language models (VLMs) typically encode substantially more visual tokens than text tokens, resulting in significant token redundancy. Pruning uninformative visual tokens is…
cs.CV2025★ 9 cited
Open-Vocabulary Camouflaged Object Segmentation with Cascaded Vision Language Models
Kai Zhao, Wubang Yuan, Zheng Wang +4
Open-Vocabulary Camouflaged Object Segmentation (OVCOS) seeks to segment and classify camouflaged objects from arbitrary categories, presenting unique challenges due to visual ambi…