11 papers
DisDop: Distillation with Domain Priors for Open-Vocabulary Aerial Object Detection
Ruihao Xu, Yong Liu, Yansong Tang +6
With the widespread application of drones in recent years, object detection of aerial images has attracted increasing attention, especially open-vocabulary aerial detection which i…
FDDet: Achieving Data-Efficient Food Defect Detection Under Real-World Scenarios
Ruihao Xu, Yong Liu, Yansong Tang
Food defect detection is critical for automated quality control, yet existing studies lack unified benchmarks and suffer from data scarcity. We introduce FDD-48, a comprehensive da…
Segment Anything with Motion, Geometry, and Semantic Adaptation for Complex Nonlinear Visual Object Tracking
Deyi Zhu, Yuji Wang, Yong Liu +4
Traditional visual object tracking (VOT) methods typically rely on task-specific supervised training, limiting their generalization to unseen objects and challenging scenarios with…
Self-Calibrated CLIP for Training-Free Open-Vocabulary Segmentation
Sule Bai, Yong Liu, Yifei Han +4
Recent advancements in pre-trained vision-language models like CLIP have enabled the task of open-vocabulary segmentation. CLIP demonstrates impressive zero-shot capabilities in va…
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
Haoji Zhang, Yiqin Wang, Yansong Tang +3
Benefiting from the advances in large language models and cross-modal alignment, existing multimodal large language models have achieved prominent performance in image and short vi…
Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation
Yong Liu, SongLi Wu, Sule Bai +3
Open-vocabulary segmentation aims to achieve segmentation of arbitrary categories given unlimited text inputs as guidance. To achieve this, recent works have focused on developing…