10 papers
SAMOT: State-Aware Step Modulation and Optimal Transport Matching for Audio-Visual Instance Segmentation
Kai Peng, Yunzhe Shen, Miao Zhang +5
Audio-Visual Instance Segmentation (AVIS) aims to simultaneously classify, segment, and track sounding objects within video sequences. Unlike Audio-Visual Semantic Segmentation (AV…
Hear to See: Discerning Stateful Listening for Audio-Visual Instance Segmentation
Leiye Liu, Miao Zhang, Jiahong Jiang +7
Audio-visual instance segmentation (AVIS) requires accurately identifying and tracking individual sounding objects with pixel-level masks. Existing methods struggle to match overla…
AdaSurvMamba: Dynamic Fusion and Semantic Scanning for Multimodal Survival Analysis
Jialong Zhong, Tingwei Liu, Baokun Yue +7
Multimodal survival analysis utilizing whole slide images (WSIs) and genomic profiles is fundamental for cancer prognosis. Recently, state-space models like Mamba have emerged as p…
Text as Illumination: Spatial Contrastive Retinex Learning for Language-guided Medical Image Segmentation
Jian Shi, Cheng Zhen, Pingping Zhang +6
Language-guided Medical Image Segmentation (LMIS) has shown great potential to improve the delineation of anatomical structures and lesions by integrating clinical textual informat…
Towards Large Model Feature Coding
Youwei Pang, Changsheng Gao, Dong Liu +2
Large models have delivered remarkable performance across a wide range of perception and generation tasks, yet practical deployment is increasingly constrained by computational and…
SAM3-I: Segment Anything with Instructions
Jingjing Li, Yue Feng, Yuchen Guo +10
Segment Anything Model 3 (SAM3) advances open-vocabulary segmentation through promptable concept segmentation, enabling users to segment all instances associated with a given conce…