works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.CV2026

Thinking in Video: Can Video Generators Really Reason About the Real World?

Yongheng Zhang, Guang Yang, Ruihan Hou +12

Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative models to simulate, predict, and reason about…

cs.AI2026

TopoAgent: A Self-Evolving Topological Agent for Multimodal Scientific Reasoning

Mingze Xu, Yinghui Li, Jiayi Kuang +5

TopoAgent introduces a graph‑based, self‑evolving framework that breaks down multimodal scientific queries into visual atoms and organizes them in a DAG, allowing dynamic refinemen…

cs.CV2026

Latent Visual Cache for Video Reasoning

Yongheng Zhang, Zhipeng Xu, Hao Wu +4

Video reasoning requires Large Multimodal Models (LMMs) to remain grounded in dense evidence, yet existing systems largely adopt "read-once, generate-many" paradigm, in which visua…

cs.AI2026

From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI

Yongheng Zhang, Ziang Liu, Jiaxuan Zhu +17

Large Language Models (LLMs) are undergoing a fundamental transformation from conversational generators into integrated AI systems capable of reasoning, action, memory, and self-im…

cs.CV2026

TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning

Daixian Liu, Jiayi Kuang, Yinghui Li +8

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in visual recognition and semantic understanding. Nevertheless, their ability to perform precise composit…

cs.SE2026

EvoConfig: Self-Evolving Multi-Agent Systems for Efficient Autonomous Environment Configuration

Xinshuai Guo, Jiayi Kuang, Linyue Pan +6

A reliable executable environment is the foundation for ensuring that large language models solve software engineering tasks. Due to the complex and tedious construction process, l…