5 papers
GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation Reasoning
Xin Xiao, Jiang Zhong, Junnan Zhu +4
Earth observation (EO) agents construct scientifically valid tool workflows and ground their conclusions in current geospatial evidence. This is challenging because EO workflows ar…
LAVE: Latent Visual Evidence-Enhanced Planning for Video Tool-use Agents
Zijian Wang, Junnan Zhu, Rongzhen Li +7
Long-video understanding requires models to efficiently acquire and reuse sparse visual evidence from long and redundant video streams. Recent video tool-use agents address this ch…
ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding
Xiao Liu, Nayu Liu, Junnan Zhu +6
Video understanding requires active evidence seeking, motivating tool-augmented video agents for temporal reasoning, cross-modal understanding, and complex question answering. Exis…
Do Models See in Line with Human Vision? Probing the Correspondence Between LVLM Representations and EEG Signals
Xin Xiao, Yang Lei, Haoyang Zeng +6
Large Vision Language Models (LVLMs) exhibit strong visual understanding and reasoning abilities. However, whether their internal representations reflect human visual cognition is…
RingMo-Agent: A Unified Remote Sensing Foundation Model for Multi-Platform and Multi-Modal Reasoning
Huiyang Hu, Peijin Wang, Yingchao Feng +9
Remote sensing (RS) images from multiple modalities and platforms exhibit diverse details due to differences in sensor characteristics and imaging perspectives. Existing vision-lan…