4 citations · 4 across the 3 of their papers we have counts for
1 paper · 1 filter
Zichen Wen, Yifeng Gao, Shaobo Wang +5
Vision tokens in multimodal large language models often dominate huge computational overhead due to their excessive length compared to linguistic modality. Abundant recent methods…