88 citations · 103 across the 10 of their papers we have counts for
15 papers
RingMoClaw: An Experience-Inspired Multi-Agent Framework for Self-Evolving Research in Remote Sensing
Kaiyue Kang, Qixuan He, Peijin Wang +9
Remote sensing visual models have continuously advanced various interpretation tasks. However, the research process behind model improvement still heavily relies on manual expertis…
GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation Reasoning
Xin Xiao, Jiang Zhong, Junnan Zhu +4
Earth observation (EO) agents construct scientifically valid tool workflows and ground their conclusions in current geospatial evidence. This is challenging because EO workflows ar…
DAPM: UAV Monocular Depth Estimation from Any Height, Pitch, Roll and FOV
Tong Ling, Wenhui Diao, Yingchao Feng +3
Monocular depth estimation is a fundamental prerequisite for 3D reconstruction and autonomous navigation in Unmanned Aerial Vehicles (UAVs). In practical deployments, UAVs operate…
Fine-grained Fragment Retrieval in Multi-modal Long-form Dialogues
Hanbo Bi, Zhiqiang Yuan, Chongyang Li +7
With the widespread adoption of multi-modal communication platforms, long-form dialogues interleaving text and images have become increasingly common. Users often need to retrieve…
CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation
Kaiwen Wei, Xiao Liu, Jie Zhang +11
Multimodal Retrieval-Augmented Generation (MRAG) enables Multimodal Large Language Models (MLLMs) to generate responses with external multimodal evidence, and numerous video-based…
ViRefSAM: Visual Reference-Guided Segment Anything Model for Remote Sensing Segmentation
Hanbo Bi, Yulong Xu, Ya Li +8
The Segment Anything Model (SAM), with its prompt-driven paradigm, exhibits strong generalization in generic segmentation tasks. However, applying SAM to remote sensing (RS) images…