2 papers
cs.AI2026
Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Pruning
Yuyao Sun, Tao Deng, Shuang Li +3
Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language tasks, but their efficiency is limited by the cost of processing numerous visual…
cs.NE2026
Do You Remember? Toward Memory-Centric Multimodal AI
Xuguang Yu, Weigang Zheng, Minyue Yu
Human memory is reconstructive, not a faithful recording. Current multimodal LLMs (MLLMs) lack this capability: they process images through a frozen visual encoder, produce a one-s…