7 papers · 1 filter
ConceptSeg-R1: Segment Any Concept via Meta-Reinforcement Learning
Yuan Zhao, Youwei Pang, Jiaming Zuo +10
Recent progress in promptable segmentation has shifted visual perception from object-level localization toward concept-level understanding. However, the notion of a concept remains…
UniPPTBench: A Unified Benchmark for Presentation Generation Across Diverse Input Settings
Bo Zhao, Maosheng Pang, Chen Zhang +3
Existing works typically focus on presentation generation under isolated input settings, whereas real-world use cases span diverse scenarios, including vague user prompts, long doc…
SAM3-I: Segment Anything with Instructions
Jingjing Li, Yue Feng, Yuchen Guo +10
Segment Anything Model 3 (SAM3) advances open-vocabulary segmentation through promptable concept segmentation, enabling users to segment all instances associated with a given conce…
Selective Noise Suppression and Discriminative Mutual Interaction for Robust Audio-Visual Segmentation
Kai Peng, Yunzhe Shen, Miao Zhang +6
The ability to capture and segment sounding objects in dynamic visual scenes is crucial for the development of Audio-Visual Segmentation (AVS) tasks. While significant progress has…
Frequency-Domain Decomposition and Recomposition for Robust Audio-Visual Segmentation
Yunzhe Shen, Kai Peng, Leiye Liu +5
Audio-visual segmentation (AVS) plays a critical role in multimodal machine learning by effectively integrating audio and visual cues to precisely segment objects or regions within…
DefMamba: Deformable Visual State Space Model
Leiye Liu, Miao Zhang, Jihao Yin +4
Recently, state space models (SSM), particularly Mamba, have attracted significant attention from scholars due to their ability to effectively balance computational efficiency and…