1 paper · 1 filter
Xuguang Yu, Weigang Zheng, Minyue Yu
Human memory is reconstructive, not a faithful recording. Current multimodal LLMs (MLLMs) lack this capability: they process images through a frozen visual encoder, produce a one-s…