2 citations · 2 across the 15 of their papers we have counts for
4 papers · 1 filter
Deep Thought Alignment: Trajectory-Level Latent Distillation for Video Reasoning
Ao Shen, Yongheng Zhang, Yinghui Li +3
Large Multimodal Models (LMMs) for video reasoning have long been hindered by the high computational cost of processing vast amounts of visual information. This dilemma motivates t…
Thinking in Video: Can Video Generators Really Reason About the Real World?
Yongheng Zhang, Guang Yang, Ruihan Hou +12
Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative models to simulate, predict, and reason about…
Latent Visual Cache for Video Reasoning
Yongheng Zhang, Zhipeng Xu, Hao Wu +4
Video reasoning requires Large Multimodal Models (LMMs) to remain grounded in dense evidence, yet existing systems largely adopt "read-once, generate-many" paradigm, in which visua…
TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning
Daixian Liu, Jiayi Kuang, Yinghui Li +8
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in visual recognition and semantic understanding, yet precise compositional spatial reasoning under geome…