3 papers
cs.CV2026
KnowVis: Knowledge-Centric Visual Summarization for Video Lectures
Yi Xu, Yifan Hou, Xiaoyu Zhang
Video lectures are valuable educational resources, but their dense and lengthy formats often overwhelm novice learners. This difficulty stems from a fundamental pedagogical mismatc…
cs.CV2026
GlanceWAM: Sparse Test-Time Imagination for World-Action Models
Linhan Wang, Zijian An, Mingyuan Zhang +7
Video generative models provide rich physical priors for robot learning, yet existing world-action models (WAMs) face a fundamental trade-off: synchronous video generation at contr…
cs.LG2026
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
Chengzu Li, Zanyi Wang, Jiaang Li +9
Vision-Language Models have excelled at textual reasoning, but they often struggle with fine-grained spatial understanding and continuous action planning, failing to simulate the d…