4 papers
ChronoVision: Temporal Reasoning via Latent State Reconstruction
Yifan Shen, Jian Xu, Boyi Li +6
Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stem…
Evaluating Cognitive Age Alignment in Interactive AI Agents
Yifan Shen, Jiawen Zhang, Jian Xu +4
While agentic AI and its core multimodal large language models (MLLMs) have demonstrated remarkable promise in language and visual reasoning across domains ranging from daily life…
LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
Chao Deng, Jiale Yuan, Pi Bu +8
Large vision language models (LVLMs) have improved the document understanding capabilities remarkably, enabling the handling of complex document elements, longer contexts, and a wi…
Multimodal Agricultural Agent Architecture (MA3): A New Paradigm for Intelligent Agricultural Decision-Making
Zhuoning Xu, Jian Xu, Mingqing Zhang +3
As a strategic pillar industry for human survival and development, modern agriculture faces dual challenges: optimizing production efficiency and achieving sustainable development.…