Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
CAM: Question Answering on Entity-Centric Videos with Continuous Extraction and Adaptive Querying
Yizhou Tian, Zizhe Chen, Shiyuan Deng +7
Memory facilitates question answering over long videos by extracting and retrieving facts to fit within the limited context windows of multimodal LLMs (MLLMs). Existing solutions t…
cs.CV2025
MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models
Garry Yang, Zizhe Chen, Man Hon Wong +5
Large Video Models (LVMs) build on the semantic capabilities of Large Language Models (LLMs) and vision modules by integrating temporal information to better understand dynamic vid…