activity
20242026
collaborators

8 papers

cs.CV2026

Routing Before Looking: Query-Adaptive Evidence Acquisition for Long-form Video Understanding

Tianyue Wang, Xuying Wu, Yuxiang Ma +7

Long-form video understanding remains challenging for video agents due to the mismatch between query demands and evidence acquisition strategies. Although recent planning-before-pe…

cs.CV2026

LUT: Latent Utility Training for Visual Reasoning

Jiaxuan Kang, Siyu Chen, Mingda Li +6

Multimodal large language models have advanced visual understanding, yet perception-intensive reasoning remains challenging. Recent latent visual reasoning methods introduce hidden…

cs.CV2026

Thinking in Video: Can Video Generators Really Reason About the Real World?

Yongheng Zhang, Guang Yang, Ruihan Hou +12

Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative models to simulate, predict, and reason about…

cs.AI2026

PRISM-MCTS: Learning from Reasoning Trajectories with Metacognitive Reflection

Siyuan Cheng, Bozhong Tian, YanChao Hao +1

PRISM-MCTS: Learning from Reasoning Trajectories with Metacognitive Reflection Siyuan Cheng, Bozhong Tian, Yanchao Hao, Zheng Wei Published: 06 Apr 2026, Last Modified: 06 Apr 2026…

cs.CL2026

Yunque DeepResearch Technical Report

Yuxuan Cai, Xinyi Lai, Peng Yuan +8

Deep research has emerged as a transformative capability for autonomous agents, empowering Large Language Models to navigate complex, open-ended tasks. However, realizing its full…

cs.CV2025

Guiding the Inner Eye: A Framework for Hierarchical and Flexible Visual Grounded Reasoning

Zhaoyang Wei, Wenchao Ding, Yanchao Hao +1

Models capable of "thinking with images" by dynamically grounding their reasoning in visual evidence represent a major leap in multimodal AI. However, replicating and advancing thi…