11 citations · 17 across the 6 of their papers we have counts for
6 papers
The Instance-centric Transformer for the RVOS Track of LSVOS Challenge: 3rd Place Solution
Bin Cao, Yisi Zhang, Hanyi Wang +2
Referring Video Object Segmentation is an emerging multi-modal task that aims to segment objects in the video given a natural language expression. In this work, we build two instan…
2nd Place Solution for MeViS Track in CVPR 2024 PVUW Workshop: Motion Expression guided Video Segmentation
Bin Cao, Yisi Zhang, Xuanxu Lin +3
Motion Expression guided Video Segmentation is a challenging task that aims at segmenting objects in the video based on natural language expressions with motion descriptions. Unlik…
Calibration & Reconstruction: Deep Integrated Language for Referring Image Segmentation
Yichen Yan, Xingjian He, Sihan Chen +1
Referring image segmentation aims to segment an object referred to by natural language expression from an image. The primary challenge lies in the efficient propagation of fine-gra…
SC-Tune: Unleashing Self-Consistent Referential Comprehension in Large Vision Language Models
Tongtian Yue, Jie Cheng, Longteng Guo +6
Recent trends in Large Vision Language Models (LVLMs) research have been increasingly focusing on advancing beyond general image understanding towards more nuanced, object-level re…
MMNet: Multi-Mask Network for Referring Image Segmentation
Yichen Yan, Xingjian He, Wenxuan Wan +1
Referring image segmentation aims to segment an object referred to by natural language expression from an image. However, this task is challenging due to the distinct data properti…
VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending
Xingjian He, Sihan Chen, Fan Ma +7
Large-scale image-text contrastive pre-training models, such as CLIP, have been demonstrated to effectively learn high-quality multimodal representations. However, there is limited…