13 citations · 13 across the 8 of their papers we have counts for
12 papers · 1 filter
SAM2Matting: Generalized Image and Video Matting
Ruiqi Shen, Guangquan Jie, Chang Liu +1
Despite impressive advances in image matting, video matting remains challenging due to the inherent gap between high-level tracking, which requires frame-wise understanding, and lo…
Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and Generation
Jinyu Liu, Xincheng Shuai, Henghui Ding +1
Unified multimodal models capable of both understanding and generation have achieved remarkable strides. However, despite their unified designs, existing evaluations typically asse…
FeVOS: Foresight Expression Video Object Segmentation
Kehan Lan, Kaining Ying, Henghui Ding
Existing Referring Video Object Segmentation tasks focus on referring expressions describing events, actions or appearances of relevant objects within the observed frames, lacking…
SAM3-DMS: Decoupled Memory Selection for Multi-target Video Segmentation of SAM3
Ruiqi Shen, Chang Liu, Henghui Ding
Segment Anything 3 (SAM3) has established a powerful foundation that robustly detects, segments, and tracks specified targets in videos. However, in its original implementation, it…
MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation
Henghui Ding, Chang Liu, Shuting He +4
This paper proposes a large-scale multi-modal dataset for referring motion expression video segmentation, focusing on segmenting and tracking target objects in videos based on lang…
Segment Anything Across Shots: A Method and Benchmark
Hengrui Hu, Kaining Ying, Henghui Ding
This work focuses on multi-shot semi-supervised video object segmentation (MVOS), which aims at segmenting the target object indicated by an initial mask throughout a video with mu…