4 papers
BindEdit: Taming Attention Leakage for Precise Multi-Object Image Editing
Chaewon Park, Soyoon Lee, Naeun Lee +3
Real image editing enables precise manipulation of visual content, yet existing methods often fail in complex multi-object scenarios, causing semantic blending, object duplication,…
CMTM: Cross-Modal Token Modulation for Unsupervised Video Object Segmentation
Inseok Jeon, Suhwan Cho, Minhyeok Lee +6
Recent advances in unsupervised video object segmentation have highlighted the potential of two-stream architectures that integrate appearance and motion cues. However, fully lever…
Cross Pseudo Labeling For Weakly Supervised Video Anomaly Detection
Dayeon Lee, Donghyeong Kim, Chaewon Park +2
Weakly supervised video anomaly detection aims to detect anomalies and identify abnormal categories with only video-level labels. We propose CPL-VAD, a dual-branch framework with c…
GenCLIP: Generalizing CLIP Prompts for Zero-shot Anomaly Detection
Donghyeong Kim, Chaewon Park, Suhwan Cho +4
Zero-shot anomaly detection (ZSAD) aims to identify anomalies in unseen categories by leveraging CLIP's zero-shot capabilities to match text prompts with visual features. A key cha…