Timeline and Boundary Guided Diffusion Network for Video Shadow Detection
arXiv:2408.11785 · doi:10.1145/3664647.3681236
Abstract
Video Shadow Detection (VSD) aims to detect the shadow masks with frame sequence. Existing works suffer from inefficient temporal learning. Moreover, few works address the VSD problem by considering the characteristic (i.e., boundary) of shadow. Motivated by this, we propose a Timeline and Boundary Guided Diffusion (TBGDiff) network for VSD where we take account of the past-future temporal guidance and boundary information jointly. In detail, we design a Dual Scale Aggregation (DSA) module for better temporal understanding by rethinking the affinity of the long-term and short-term frames for the clipped video. Next, we introduce Shadow Boundary Aware Attention (SBAA) to utilize the edge contexts for capturing the characteristics of shadows. Moreover, we are the first to introduce the Diffusion model for VSD in which we explore a Space-Time Encoded Embedding (STEE) to inject the temporal guidance for Diffusion to conduct shadow detection. Benefiting from these designs, our model can not only capture the temporal information but also the shadow property. Extensive experiments show that the performance of our approach overtakes the state-of-the-art methods, verifying the effectiveness of our components. We release the codes, weights, and results at \url{https://github.com/haipengzhou856/TBGDiff}.
ACM MM2024
References in corpus (11)
- SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object Segmentation
- Revisiting Shadow Detection: A New Benchmark Dataset for Complex World
- Learning from Synthetic Shadows for Shadow Detection and Removal
- Video Object Segmentation with Adaptive Feature Bank and Uncertain-Region Refinement
- HybridMIM: A Hybrid Masked Image Modeling Framework for 3D Medical Image Segmentation
- RainMamba: Enhanced Locality Learning with State Space Models for Video Deraining
- CARD: Classification and Regression Diffusion Models
- Dynamic Interactive Relation Capturing via Scene Graph Learning for Robotic Surgical Report Generation
- Online Unsupervised Video Object Segmentation via Contrastive Motion Clustering
- Language-Driven Interactive Shadow Detection