2 papers
cs.CV2026
Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents
Wenqi Liu, Shijie Ma, Yunxiao Wang +21
Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While…
cs.CL2026
Knowing Isn't Always Saying: When Do Spatial Encodings Reach Answers in Vision-Language Models?
Zeyu Wang, Xinming Xu
Vision-language models are known to encode spatial information in their hidden states, yet often fail to use it when answering. However, it remains unclear when and where this enco…