5 papers
Understand Before Detect: Vision--Language Learning for Omni-Domain Infrared Small Target Detection
Haoyang Yuan, Boyang Li, Yingqian Wang +7
Omni-domain infrared small target (IRST) detection is crucial for infrared surveillance, yet remains challenging due to heterogeneous imaging domains and inconsistent target charac…
Focus on What Really Matters in Low-Altitude Governance: A Management-Centric Multi-Modal Benchmark with Implicitly Coordinated Vision-Language Reasoning Framework
Hao Chang, Zhihui Wang, Lingxiang Wu +5
Low-altitude vision systems are becoming a critical infrastructure for smart city governance. However, existing object-centric perception paradigms and loosely coupled vision-langu…
Visible-Thermal Tiny Object Detection: A Benchmark Dataset and Baselines
Xinyi Ying, Chao Xiao, Ruojing Li +13
Small object detection (SOD) has been a longstanding yet challenging task for decades, with numerous datasets and algorithms being developed. However, they mainly focus on either v…
Heterogeneous Graph Transformer for Multiple Tiny Object Tracking in RGB-T Videos
Qingyu Xu, Longguang Wang, Weidong Sheng +4
Tracking multiple tiny objects is highly challenging due to their weak appearance and limited features. Existing multi-object tracking algorithms generally focus on single-modality…
Highly Efficient and Unsupervised Framework for Moving Object Detection in Satellite Videos
C. Xiao, W. An, Y. Zhang +5
Moving object detection in satellite videos (SVMOD) is a challenging task due to the extremely dim and small target characteristics. Current learning-based methods extract spatio-t…