4 papers
RingMoClaw: An Experience-Inspired Multi-Agent Framework for Self-Evolving Research in Remote Sensing
Kaiyue Kang, Qixuan He, Peijin Wang +9
Remote sensing visual models have continuously advanced various interpretation tasks. However, the research process behind model improvement still heavily relies on manual expertis…
GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation Reasoning
Xin Xiao, Jiang Zhong, Junnan Zhu +4
Earth observation (EO) agents construct scientifically valid tool workflows and ground their conclusions in current geospatial evidence. This is challenging because EO workflows ar…
LAVE: Latent Visual Evidence-Enhanced Planning for Video Tool-use Agents
Zijian Wang, Junnan Zhu, Rongzhen Li +7
Long-video understanding requires models to efficiently acquire and reuse sparse visual evidence from long and redundant video streams. Recent video tool-use agents address this ch…
Do Models See in Line with Human Vision? Probing the Correspondence Between LVLM Representations and EEG Signals
Xin Xiao, Yang Lei, Haoyang Zeng +6
Large Vision Language Models (LVLMs) exhibit strong visual understanding and reasoning abilities. However, whether their internal representations reflect human visual cognition is…