4 papers
IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning
Zhichao Sun, Yidong Ma, Gang Liu +4
Large Vision-Language Models (LVLMs) achieve impressive performance across multiple tasks. A significant challenge, however, is their prohibitive inference cost when processing hig…
CQ-DINO: Mitigating Gradient Dilution via Category Queries for Vast Vocabulary Object Detection
Zhichao Sun, Huazhang Hu, Yidong Ma +5
With the exponential growth of data, traditional object detection methods are increasingly struggling to handle vast vocabulary object detection tasks effectively. We analyze two k…
RefDrone: A Challenging Benchmark for Referring Expression Comprehension in Drone Scenes
Zhichao Sun, Yepeng Liu, Zhiling Su +7
Drones have become prevalent robotic platforms with diverse applications, showing significant potential in Embodied Artificial Intelligence (Embodied AI). Referring Expression Comp…
MIFNet: Learning Modality-Invariant Features for Generalizable Multimodal Image Matching
Yepeng Liu, Zhichao Sun, Baosheng Yu +4
Many keypoint detection and description methods have been proposed for image matching or registration. While these methods demonstrate promising performance for single-modality ima…