256 citations · 437 across the 37 of their papers we have counts for
1 paper · 1 filter
Xiaoran Fan, Tao Ji, Changhao Jiang +21
Current large vision-language models (VLMs) often encounter challenges such as insufficient capabilities of a single visual component and excessively long visual tokens. These issu…