5 papers · 1 filter
SAM 3: Segment Anything with Concepts
Nicolas Carion, Laura Gustafson, Yuan-Ting Hu +35
We present Segment Anything Model (SAM) 3, a unified model that detects, segments, and tracks objects in images and videos based on concept prompts, which we define as either short…
Calibrating Undisciplined Over-Smoothing in Transformer for Weakly Supervised Semantic Segmentation
Lechao Cheng, Zerun Liu, Jingxuan He +3
Weakly supervised semantic segmentation (WSSS) has recently attracted considerable attention because it requires fewer annotations than fully supervised approaches, making it espec…
AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection
Bin-Bin Gao, Yue Zhou, Jiangtao Yan +7
Universal visual anomaly detection aims to identify anomalies from novel or unseen vision domains without additional fine-tuning, which is critical in open scenarios. Recent studie…
Rethinking Visual Content Refinement in Low-Shot CLIP Adaptation
Jinda Lu, Shuo Wang, Yanbin Hao +3
Recent adaptations can boost the low-shot capability of Contrastive Vision-Language Pre-training (CLIP) by effectively facilitating knowledge transfer. However, these adaptation me…
PosMLP-Video: Spatial and Temporal Relative Position Encoding for Efficient Video Recognition
Yanbin Hao, Diansong Zhou, Zhicai Wang +2
In recent years, vision Transformers and MLPs have demonstrated remarkable performance in image understanding tasks. However, their inherently dense computational operators, such a…