1 paper · 1 filter
Bin Kang, Bin Chen, Junjie Wang +3
Existing Visual Language Models (VLMs) suffer structural limitations where a few low contribution tokens may excessively capture global semantics, dominating the information aggreg…