1 paper
Bin Kang, Bin Chen, Junjie Wang +3
Existing Visual Language Models (VLMs) suffer structural limitations where a few low contribution tokens may excessively capture global semantics, dominating the information aggreg…