2 citations · 4 across the 4 of their papers we have counts for
4 papers
VISA: Reasoning Video Object Segmentation via Large Language Models
Cilin Yan, Haochen Wang, Shilin Yan +5
Existing Video Object Segmentation (VOS) relies on explicit user instructions, such as categories, masks, or short phrases, restricting their ability to perform complex video segme…
Mining Open Semantics from CLIP: A Relation Transition Perspective for Few-Shot Learning
Cilin Yan, Haochen Wang, Xiaolong Jiang +4
Contrastive Vision-Language Pre-training(CLIP) demonstrates impressive zero-shot capability. The key to improve the adaptation of CLIP to downstream task with few exemplars lies in…
RSBuilding: Towards General Remote Sensing Image Building Extraction and Change Detection with Foundation Model
Mingze Wang, Lili Su, Cilin Yan +4
The intelligent interpretation of buildings plays a significant role in urban planning and management, macroeconomic analysis, population dynamics, etc. Remote sensing image buildi…
1st Place Solution for the 5th LSVOS Challenge: Video Instance Segmentation
Tao Zhang, Xingye Tian, Yikang Zhou +7
Video instance segmentation is a challenging task that serves as the cornerstone of numerous downstream applications, including video editing and autonomous driving. In this report…