1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Yichen Guo, Hanze Li, Zonghao Zhang +3
Although large vision-language models (LVLMs) leverage rich visual token representations to achieve strong performance on multimodal tasks, these tokens also introduce significant…