2 papers
cs.CV2026
CausalChapter: Improving Long-Video Chaptering with Interventional Dependency Modeling
Xinran Duan, Guozhang Li, Yaoyao Zhong +3
Long-form instructional videos require automatic chaptering to support browsing, navigation, and knowledge access. Recent long-context language models can perform chaptering from t…
cs.CV2026
Who Drives the Probability Game of VLMs? A Temporal Causal Drive Evaluation Framework
Shuyao Xiao, Shengling Wang, Haoyu Niu +4
Vision-language models (VLMs) are increasingly evaluated on complex image and video understanding tasks, yet conventional metrics primarily assess final-answer quality and reveal l…