19 papers
Skeleton-to-Image Encoding: Enabling Skeleton Representation Learning via Vision-Pretrained Models
Siyuan Yang, Jun Liu, Hao Cheng +5
Recent advances in large-scale pretrained vision models have demonstrated impressive capabilities across a wide range of downstream tasks, including cross-modal and multi-modal sce…
Cross-Domain Few-Shot Segmentation via Multi-view Progressive Adaptation
Jiahao Nie, Guanqiao Fu, Wenbin An +3
Cross-Domain Few-Shot Segmentation aims to segment categories in data-scarce domains conditioned on a few exemplars. Typical methods first establish few-shot capability in a large-…
Boosting SAM for Cross-Domain Few-Shot Segmentation via Conditional Point Sparsification
Jiahao Nie, Yun Xing, Wenbin An +6
Motivated by the success of the Segment Anything Model (SAM) in promptable segmentation, recent studies leverage SAM to develop training-free solutions for few-shot segmentation, w…
E.M.Ground: A Temporal Grounding Vid-LLM with Holistic Event Perception and Matching
Jiahao Nie, Wenbin An, Gongjie Zhang +4
Despite recent advances in Video Large Language Models (Vid-LLMs), Temporal Video Grounding (TVG), which aims to precisely localize time segments corresponding to query events, rem…
MMRel: Benchmarking Relation Understanding in Multi-Modal Large Language Models
Jiahao Nie, Gongjie Zhang, Wenbin An +4
Though Multi-modal Large Language Models (MLLMs) have recently achieved significant progress, they often struggle to understand diverse and complicated inter-object relations. Spec…
Class-Independent Increment: An Efficient Approach for Multi-label Class-Incremental Learning
Chenhao Ding, Songlin Dong, Zhengdong Zhou +4
Current research on class-incremental learning primarily focuses on single-label classification tasks. However, real-world applications often involve multi-label scenarios, such as…