most citedVISA: Reasoning Video Object Segmentation via Large Language Models

2 citations · 3 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV2024

Attention Head Purification: A New Perspective to Harness CLIP for Domain Generalization

Yingfan Wang, Guoliang Kang

Domain Generalization (DG) aims to learn a model from multiple source domains to achieve satisfactory performance on unseen target domains. Recent works introduce CLIP to DG tasks…

cs.CV2024

SLCA++: Unleash the Power of Sequential Fine-tuning for Continual Learning with Pre-training

Gengwei Zhang, Liyuan Wang, Guoliang Kang +2

In recent years, continual learning with pre-training (CLPT) has received widespread interest, instead of its traditional focus of training from scratch. The use of strong pre-trai…

cs.CV20242 cited

VISA: Reasoning Video Object Segmentation via Large Language Models

Cilin Yan, Haochen Wang, Shilin Yan +5

Existing Video Object Segmentation (VOS) relies on explicit user instructions, such as categories, masks, or short phrases, restricting their ability to perform complex video segme…

cs.CV2024

Mining Open Semantics from CLIP: A Relation Transition Perspective for Few-Shot Learning

Cilin Yan, Haochen Wang, Xiaolong Jiang +4

Contrastive Vision-Language Pre-training(CLIP) demonstrates impressive zero-shot capability. The key to improve the adaptation of CLIP to downstream task with few exemplars lies in…

cs.CV20231 cited

LatentWarp: Consistent Diffusion Latents for Zero-Shot Video-to-Video Translation

Yuxiang Bao, Di Qiu, Guoliang Kang +4

Leveraging the generative ability of image diffusion models offers great potential for zero-shot video-to-video translation. The key lies in how to maintain temporal consistency ac…