5 papers · 1 filter
Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models
Yi Ding, Lijun Li, Bing Cao +1
Large Vision-Language Models (VLMs) have achieved remarkable performance across a wide range of tasks. However, their deployment in safety-critical domains poses significant challe…
LSVOS Challenge Report: Large-scale Complex and Long Video Object Segmentation
Henghui Ding, Lingyi Hong, Chang Liu +30
Despite the promising performance of current video segmentation models on existing benchmarks, these models still struggle with complex scenes. In this paper, we introduce the 6th…
The Instance-centric Transformer for the RVOS Track of LSVOS Challenge: 3rd Place Solution
Bin Cao, Yisi Zhang, Hanyi Wang +2
Referring Video Object Segmentation is an emerging multi-modal task that aims to segment objects in the video given a natural language expression. In this work, we build two instan…
PVUW 2024 Challenge on Complex Video Understanding: Methods and Results
Henghui Ding, Chang Liu, Yunchao Wei +34
Pixel-level Video Understanding in the Wild Challenge (PVUW) focus on complex video understanding. In this CVPR 2024 workshop, we add two new tracks, Complex Video Object Segmentat…
2nd Place Solution for MeViS Track in CVPR 2024 PVUW Workshop: Motion Expression guided Video Segmentation
Bin Cao, Yisi Zhang, Xuanxu Lin +3
Motion Expression guided Video Segmentation is a challenging task that aims at segmenting objects in the video based on natural language expressions with motion descriptions. Unlik…