collaborators

5 papers

cs.CV2026

Training-Free Open-Vocabulary Visual Grounding for Remote Sensing Images and Videos

Ke Li, Di Wang, Yongshan Zhu +5

Remote sensing visual grounding (RSVG) aims to localize a referred target in a remote sensing image or video according to a natural language expression. Existing RSVG methods usual…

cs.CV2026

ProVG: Progressive Visual Grounding via Language Decoupling for Remote Sensing Imagery

Ke Li, Ting Wang, Di Wang +4

Remote sensing visual grounding (RSVG) aims to localize objects in remote sensing imagery according to natural language expressions. Previous methods typically rely on sentence-lev…

cs.CV2025

Scale-Aware Relay and Scale-Adaptive Loss for Tiny Object Detection in Aerial Images

Jinfu Li, Yuqi Huang, Hong Song +5

Recently, despite the remarkable advancements in object detection, modern detectors still struggle to detect tiny objects in aerial images. One key reason is that tiny objects carr…

cs.CV2025

RSVG-ZeroOV: Exploring a Training-Free Framework for Zero-Shot Open-Vocabulary Visual Grounding in Remote Sensing Images

Ke Li, Di Wang, Ting Wang +6

Remote sensing visual grounding (RSVG) aims to localize objects in remote sensing images based on free-form natural language expressions. Existing approaches are typically constrai…

cs.CV2025

ReWind: Understanding Long Videos with Instructed Learnable Memory

Anxhelo Diko, Tinghuai Wang, Wassim Swaileh +2

Vision-Language Models (VLMs) are crucial for applications requiring integrated understanding textual and visual information. However, existing VLMs struggle with long videos due t…