1 paper · 1 filter
Yizhou Tian, Zizhe Chen, Shiyuan Deng +7
Memory facilitates question answering over long videos by extracting and retrieving facts to fit within the limited context windows of multimodal LLMs (MLLMs). Existing solutions t…