From the 1 of 10 linked papers with an AI index.
10 papers
LUT: Latent Utility Training for Visual Reasoning
Jiaxuan Kang, Siyu Chen, Mingda Li +6
Multimodal large language models have advanced visual understanding, yet perception-intensive reasoning remains challenging. Recent latent visual reasoning methods introduce hidden…
See2Think: Do Multimodal Models Really Use Intermediate Visual States?
Siyu Yan, Zhuoran Yan, Haiying Xu +10
The paper presents See2Think, an evaluation framework and benchmark for testing whether multimodal large language models actually use intermediate visual states during reasoning, a…
Thinking in Video: Can Video Generators Really Reason About the Real World?
Yongheng Zhang, Guang Yang, Ruihan Hou +12
Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative models to simulate, predict, and reason about…
Latent Visual Cache for Video Reasoning
Yongheng Zhang, Zhipeng Xu, Hao Wu +4
Video reasoning requires Large Multimodal Models (LMMs) to remain grounded in dense evidence, yet existing systems largely adopt "read-once, generate-many" paradigm, in which visua…
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
Yongheng Zhang, Ziang Liu, Jiaxuan Zhu +17
Large Language Models (LLMs) are undergoing a fundamental transformation from conversational generators into integrated AI systems capable of reasoning, action, memory, and self-im…
Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding
Yinghui Li, Jiayi Kuang, Peng Xing +11
Multimodal large language models (MLLMs) perform strongly on natural images, yet their ability to understand discrete visual symbols remains unclear. We present a multi-domain benc…