9 papers
NTIRE 2026 Challenge on Video Saliency Prediction: Methods and Results
Andrey Moskalenko, Alexey Bryncev, Ivan Kosmynin +40
This paper presents an overview of the NTIRE 2026 Challenge on Video Saliency Prediction. The goal of the challenge participants was to develop automatic saliency map prediction me…
Camouflage-aware Image-Text Retrieval via Expert Collaboration
Yao Jiang, Zhongkuan Mao, Xuan Wu +2
Camouflaged scene understanding (CSU) has attracted significant attention due to its broad practical implications. However, in this field, robust image-text cross-modal alignment r…
Samba+: General and Accurate Salient Object Detection via A More Unified Mamba-based Framework
Wenzhuo Zhao, Keren Fu, Jiahao He +3
Existing salient object detection (SOD) models are generally constrained by the limited receptive fields of convolutional neural networks (CNNs) and quadratic computational complex…
MoGen: A Unified Collaborative Framework for Controllable Multi-Object Image Generation
Yanfeng Li, Yue Sun, Keren Fu +5
Existing multi-object image generation methods face difficulties in achieving precise alignment between localized image generation regions and their corresponding semantics based o…
Unleashing the Power of Motion and Depth: A Selective Fusion Strategy for RGB-D Video Salient Object Detection
Jiahao He, Daerji Suolang, Keren Fu +1
Applying salient object detection (SOD) to RGB-D videos is an emerging task called RGB-D VSOD and has recently gained increasing interest, due to considerable performance gains of…
Mamba-based Spatio-Frequency Motion Perception for Video Camouflaged Object Detection
Xin Li, Keren Fu, Qijun Zhao
Existing video camouflaged object detection (VCOD) methods primarily rely on spatial appearances for motion perception. However, the high foreground-background similarity in VCOD l…