3 papers
cs.CV2026
Plug, Play, and Fortify: A Low-Cost Module for Robust Multimodal Image Understanding Models
Siqi Lu, Wanying Xu, Yongbin Zheng +3
Missing modalities present a fundamental challenge in multimodal models, often causing catastrophic performance degradation. Our observations suggest that this fragility stems from…
cs.CV2025
VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection
Jianhang Yao, Yongbin Zheng, Siqi Lu +2
To identify objects beyond predefined categories, open-vocabulary aerial object detection (OVAD) leverages the zero-shot capabilities of visual-language models (VLMs) to generalize…
cs.CV2023
Metric-aligned Sample Selection and Critical Feature Sampling for Oriented Object Detection
Peng Sun, Yongbin Zheng, Wenqi Wu +2
Arbitrary-oriented object detection is a relatively emerging but challenging task. Although remarkable progress has been made, there still remain many unsolved issues due to the la…