collaborators

10 papers

cs.CV2026

Training-Free Open-Vocabulary Visual Grounding for Remote Sensing Images and Videos

Ke Li, Di Wang, Yongshan Zhu +5

Remote sensing visual grounding (RSVG) aims to localize a referred target in a remote sensing image or video according to a natural language expression. Existing RSVG methods usual…

cs.CV2026

OmniTraffic: A Controllable Generation Pipeline and Benchmark for Spatio-Temporal Traffic Reasoning

Maonan Wang, Zhengyan Huang, Kemou Jiang +13

Traffic scene understanding requires models to reason beyond object recognition, including lane topology, multi-view geometry, temporal evolution, and signal-phase semantics. Howev…

cs.RO2026

Y-BotFrame: An Extensible Embodied Agent Framework for Quadruped Robot Assistants

Luyao Zhang, Ke Li, Yuan Ding +9

Quadruped robots are capable of traversing a wide range of complex terrains with high flexibility. As highly mobile ground-based intelligent platforms, they can be equipped with mo…

cs.RO2026

AerialClaw: An Open-Source Framework for LLM-Driven Autonomous Aerial Agents

Ke Li, Jianfei Yang, Luyao Zhang +8

Unmanned aerial vehicles (UAVs) are increasingly used in inspection, search and rescue, environmental monitoring, and emergency response. However, most UAV applications still rely…

cs.CV2026

ProVG: Progressive Visual Grounding via Language Decoupling for Remote Sensing Imagery

Ke Li, Ting Wang, Di Wang +4

Remote sensing visual grounding (RSVG) aims to localize objects in remote sensing imagery according to natural language expressions. Previous methods typically rely on sentence-lev…

cs.CV2025

Improving Few-Shot Change Detection Visual Question Answering via Decision-Ambiguity-guided Reinforcement Fine-Tuning

Fuyu Dong, Ke Li, Di Wang +5

Change detection visual question answering (CDVQA) requires answering text queries by reasoning about semantic changes in bi-temporal remote sensing images. A straightforward appro…