14 papers
SAM-MT: Real-Time Interactive Multi-Target Video Segmentation
Ruiqi Shen, Chang Liu, Henghui Ding
Modern Video Object Segmentation (VOS) involves tracking and segmenting user-specified targets. While recent approaches have achieved remarkable performance in single-target scenar…
SAM2Matting: Generalized Image and Video Matting
Ruiqi Shen, Guangquan Jie, Chang Liu +1
Despite impressive advances in image matting, video matting remains challenging due to the inherent gap between high-level tracking, which requires frame-wise understanding, and lo…
Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and Generation
Jinyu Liu, Xincheng Shuai, Henghui Ding +1
Unified multimodal models capable of both understanding and generation have achieved remarkable strides. However, despite their unified designs, existing evaluations typically asse…
FeVOS: Foresight Expression Video Object Segmentation
Kehan Lan, Kaining Ying, Henghui Ding
Existing Referring Video Object Segmentation tasks focus on referring expressions describing events, actions or appearances of relevant objects within the observed frames, lacking…
AVTrack: Audio-Visual Tracking in Human-centric Complex Scenes
Yaoting Wang, Yun Zhou, Zipei Zhang +1
Audio-visual speaker tracking aims to localize and track active speakers by leveraging auditory and visual cues, enabling fine-grained, human-centric scene understanding. This capa…
SAM3-DMS: Decoupled Memory Selection for Multi-target Video Segmentation of SAM3
Ruiqi Shen, Chang Liu, Henghui Ding
Segment Anything 3 (SAM3) has established a powerful foundation that robustly detects, segments, and tracks specified targets in videos. However, in its original implementation, it…