5 papers
IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning
Zhichao Sun, Yidong Ma, Gang Liu +4
Large Vision-Language Models (LVLMs) achieve impressive performance across multiple tasks. A significant challenge, however, is their prohibitive inference cost when processing hig…
RefDrone: A Challenging Benchmark for Referring Expression Comprehension in Drone Scenes
Zhichao Sun, Yepeng Liu, Zhiling Su +7
Drones have become prevalent robotic platforms with diverse applications, showing significant potential in Embodied Artificial Intelligence (Embodied AI). Referring Expression Comp…
CQ-DINO: Mitigating Gradient Dilution via Category Queries for Vast Vocabulary Object Detection
Zhichao Sun, Huazhang Hu, Yidong Ma +5
With the exponential growth of data, traditional object detection methods are increasingly struggling to handle vast vocabulary object detection tasks effectively. We analyze two k…
MIFNet: Learning Modality-Invariant Features for Generalizable Multimodal Image Matching
Yepeng Liu, Zhichao Sun, Baosheng Yu +4
Many keypoint detection and description methods have been proposed for image matching or registration. While these methods demonstrate promising performance for single-modality ima…
Shape Transformation Driven by Active Contour for Class-Imbalanced Semi-Supervised Medical Image Segmentation
Yuliang Gu, Yepeng Liu, Zhichao Sun +3
Annotating 3D medical images demands expert knowledge and is time-consuming. As a result, semi-supervised learning (SSL) approaches have gained significant interest in 3D medical i…