1 paper
Yijun Chen, Yaqi Zheng, Yanya Li +11
Multimodal memory offers a scalable interface for long-video question answering, but existing methods often retrieve captions, frames, transcripts, summaries, or graph facts as iso…