9 papers · 1 filter
GA2-CLIP: Generic Attribute Anchor for Efficient Prompt Tuningin Video-Language Models
Bin Wang, Ruotong Hu, Wentong Li +5
Visual and textual soft prompt tuning can effectively improve the adaptability of Vision-Language Models (VLMs) in downstream tasks. However, fine-tuning on video tasks impairs the…
Divide-and-Conquer Decoupled Network for Cross-Domain Few-Shot Segmentation
Runmin Cong, Anpeng Wang, Bin Wan +3
Cross-domain few-shot segmentation (CD-FSS) aims to tackle the dual challenge of recognizing novel classes and adapting to unseen domains with limited annotations. However, encoder…
The 1st Solution for 4th PVUW MeViS Challenge: Unleashing the Potential of Large Multimodal Models for Referring Video Segmentation
Hao Fang, Runmin Cong, Xiankai Lu +2
Motion expression video segmentation is designed to segment objects in accordance with the input motion expressions. In contrast to the conventional Referring Video Object Segmenta…
Query-guided Prototype Evolution Network for Few-Shot Segmentation
Runmin Cong, Hang Xiong, Jinpeng Chen +3
Previous Few-Shot Segmentation (FSS) approaches exclusively utilize support features for prototype generation, neglecting the specific requirements of the query. To address this, w…
SDDNet: Style-guided Dual-layer Disentanglement Network for Shadow Detection
Runmin Cong, Yuchen Guan, Jinpeng Chen +3
Despite significant progress in shadow detection, current methods still struggle with the adverse impact of background color, which may lead to errors when shadows are present on c…
Point-aware Interaction and CNN-induced Refinement Network for RGB-D Salient Object Detection
Runmin Cong, Hongyu Liu, Chen Zhang +4
By integrating complementary information from RGB image and depth map, the ability of salient object detection (SOD) for complex and challenging scenes can be improved. In recent y…