20 citations · 20 across the 14 of their papers we have counts for
14 papers
Deep Thought Alignment: Trajectory-Level Latent Distillation for Video Reasoning
Ao Shen, Yongheng Zhang, Yinghui Li +3
Large Multimodal Models (LMMs) for video reasoning have long been hindered by the high computational cost of processing vast amounts of visual information. This dilemma motivates t…
ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval
Chunyi Peng, Zhipeng Xu, Yukun Yan +9
Visual document retrieval is a critical component of multimodal retrieval-augmented generation, aiming to identify query-relevant pages from document collections where evidence is…
LUT: Latent Utility Training for Visual Reasoning
Jiaxuan Kang, Siyu Chen, Mingda Li +6
Multimodal large language models have advanced visual understanding, yet perception-intensive reasoning remains challenging. Recent latent visual reasoning methods introduce hidden…
See2Think: Do Multimodal Models Really Use Intermediate Visual States?
Siyu Yan, Zhuoran Yan, Haiying Xu +10
Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these vi…
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
Yongheng Zhang, Ziang Liu, Jiaxuan Zhu +17
Large Language Models (LLMs) are undergoing a fundamental transformation from conversational generators into integrated AI systems capable of reasoning, action, memory, and self-im…
Thinking in Video: Can Video Generators Really Reason About the Real World?
Yongheng Zhang, Guang Yang, Ruihan Hou +12
Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative models to simulate, predict, and reason about…