Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
CAM: Question Answering on Entity-Centric Videos with Continuous Extraction and Adaptive Querying
Yizhou Tian, Zizhe Chen, Shiyuan Deng +7
Memory facilitates question answering over long videos by extracting and retrieving facts to fit within the limited context windows of multimodal LLMs (MLLMs). Existing solutions t…
cs.CV2026
VISD: Enhancing Video Reasoning via Structured Self-Distillation
Hao Lin, Kunyang Lv, Xu Jiang +5
Training VideoLLMs for complex reasoning remains challenging due to sparse sequence level rewards and the lack of fine grained credit assignment over long, temporally grounded reas…