From the 1 of 16 linked papers with an AI index.
16 papers
Thinking in Video: Can Video Generators Really Reason About the Real World?
Yongheng Zhang, Guang Yang, Ruihan Hou +12
Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative models to simulate, predict, and reason about…
TopoAgent: A Self-Evolving Topological Agent for Multimodal Scientific Reasoning
Mingze Xu, Yinghui Li, Jiayi Kuang +5
TopoAgent introduces a graph‑based, self‑evolving framework that breaks down multimodal scientific queries into visual atoms and organizes them in a DAG, allowing dynamic refinemen…
Latent Visual Cache for Video Reasoning
Yongheng Zhang, Zhipeng Xu, Hao Wu +4
Video reasoning requires Large Multimodal Models (LMMs) to remain grounded in dense evidence, yet existing systems largely adopt "read-once, generate-many" paradigm, in which visua…
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
Yongheng Zhang, Ziang Liu, Jiaxuan Zhu +17
Large Language Models (LLMs) are undergoing a fundamental transformation from conversational generators into integrated AI systems capable of reasoning, action, memory, and self-im…
MoG: Mixture of Experts for Graph-based Retrieval-Augmented Generation
Zheng Yuan, Chuang Zhou, Linhao Luo +4
Retrieval-augmented generation is intensively studied to ground large language models on external evidence. However, retrieving from a unified knowledge base could inevitably intro…
Deep Tabular Research via Continual Experience-Driven Execution
Junnan Dong, Chuang Zhou, Zheng Yuan +7
Large language models often struggle with complex long-horizon analytical tasks over unstructured tables, which typically feature hierarchical and bidirectional headers and non-can…