2 papers
cs.CV2026
CausalChapter: Improving Long-Video Chaptering with Interventional Dependency Modeling
Xinran Duan, Guozhang Li, Yaoyao Zhong +3
Long-form instructional videos require automatic chaptering to support browsing, navigation, and knowledge access. Recent long-context language models can perform chaptering from t…
cs.CV2026
SketchJudge: A Diagnostic Benchmark for Grading Hand-drawn Diagrams with Multimodal Large Language Models
Yuhang Su, Mei Wang, Yaoyao Zhong +4
While Multimodal Large Language Models (MLLMs) have achieved remarkable progress in visual understanding, they often struggle when faced with the unstructured and ambiguous nature…