8 papers
Routing Before Looking: Query-Adaptive Evidence Acquisition for Long-form Video Understanding
Tianyue Wang, Xuying Wu, Yuxiang Ma +7
Long-form video understanding remains challenging for video agents due to the mismatch between query demands and evidence acquisition strategies. Although recent planning-before-pe…
LUT: Latent Utility Training for Visual Reasoning
Jiaxuan Kang, Siyu Chen, Mingda Li +6
Multimodal large language models have advanced visual understanding, yet perception-intensive reasoning remains challenging. Recent latent visual reasoning methods introduce hidden…
Thinking in Video: Can Video Generators Really Reason About the Real World?
Yongheng Zhang, Guang Yang, Ruihan Hou +12
Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative models to simulate, predict, and reason about…
PRISM-MCTS: Learning from Reasoning Trajectories with Metacognitive Reflection
Siyuan Cheng, Bozhong Tian, YanChao Hao +1
PRISM-MCTS: Learning from Reasoning Trajectories with Metacognitive Reflection Siyuan Cheng, Bozhong Tian, Yanchao Hao, Zheng Wei Published: 06 Apr 2026, Last Modified: 06 Apr 2026…
Yunque DeepResearch Technical Report
Yuxuan Cai, Xinyi Lai, Peng Yuan +8
Deep research has emerged as a transformative capability for autonomous agents, empowering Large Language Models to navigate complex, open-ended tasks. However, realizing its full…
Guiding the Inner Eye: A Framework for Hierarchical and Flexible Visual Grounded Reasoning
Zhaoyang Wei, Wenchao Ding, Yanchao Hao +1
Models capable of "thinking with images" by dynamically grounding their reasoning in visual evidence represent a major leap in multimodal AI. However, replicating and advancing thi…