1 paper
Yizhou Tian, Zizhe Chen, Shiyuan Deng +7
Memory facilitates question answering over long videos by extracting and retrieving facts to fit within the limited context windows of multimodal LLMs (MLLMs). Existing solutions t…