collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

Video-HolmesV2: Can MLLMs Reason with Spatio-Temporal Audio-Visual Evidence in Long Videos?

Zhaoyang Wei, Zipeng Wang, Yushe Cao +10

Multimodal Large Language Models have demonstrated impressive video understanding, yet their ability to reason over long-form narratives is often masked by visual-centric evaluatio…

cs.CV2026

Routing Before Looking: Query-Adaptive Evidence Acquisition for Long-form Video Understanding

Tianyue Wang, Xuying Wu, Yuxiang Ma +7

Long-form video understanding remains challenging for video agents due to the mismatch between query demands and evidence acquisition strategies. Although recent planning-before-pe…

cs.CV2026

LUT: Latent Utility Training for Visual Reasoning

Jiaxuan Kang, Siyu Chen, Mingda Li +6

Multimodal large language models have advanced visual understanding, yet perception-intensive reasoning remains challenging. Recent latent visual reasoning methods introduce hidden…

cs.CV2026

Thinking in Video: Can Video Generators Really Reason About the Real World?

Yongheng Zhang, Guang Yang, Ruihan Hou +12

Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative models to simulate, predict, and reason about…

cs.CV2026

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization

Tao Cheng, Shi-Zhe Chen, Hao Zhang +3

Chain-of-Thought (CoT) reasoning significantly elevates the complex problem-solving capabilities of multimodal large language models (MLLMs). However, adapting CoT to vision typica…