Progressively Normalized Self-Attention Network for Video Polyp Segmentation
arXiv:2105.08468 · doi:10.1007/978-3-030-87193-2_14
Abstract
Existing video polyp segmentation (VPS) models typically employ convolutional neural networks (CNNs) to extract features. However, due to their limited receptive fields, CNNs can not fully exploit the global temporal and spatial information in successive video frames, resulting in false-positive segmentation results. In this paper, we propose the novel PNS-Net (Progressively Normalized Self-attention Network), which can efficiently learn representations from polyp videos with real-time speed (~140fps) on a single RTX 2080 GPU and no post-processing. Our PNS-Net is based solely on a basic normalized self-attention block, equipping with recurrence and CNNs entirely. Experiments on challenging VPS datasets demonstrate that the proposed PNS-Net achieves state-of-the-art performance. We also conduct extensive experiments to study the effectiveness of the channel split, soft-attention, and progressive learning strategy. We find that our PNS-Net works well under different settings, making it a promising solution to the VPS task.
MICCAI 2021 (Provisional accept); Code: https://github.com/GewelsJI/PNS-Net
References in corpus (3)
Cited by in corpus (11)
- Polyp-PVT: Polyp Segmentation with Pyramid Vision Transformers
- Deep Gradient Learning for Efficient Camouflaged Object Detection
- Fast Camouflaged Object Detection via Edge-based Reversible Re-calibration Network
- ZoomNeXt: A Unified Collaborative Pyramid Network for Camouflaged Object Detection
- Full-Duplex Strategy for Video Object Segmentation
- Trichomonas Vaginalis Segmentation in Microscope Images
- Rethinking Polyp Segmentation from an Out-of-Distribution Perspective
- Frontiers in Intelligent Colonoscopy
- Depth Quality-Inspired Feature Manipulation for Efficient RGB-D Salient Object Detection
- Segmentation of kidney stones in endoscopic video feeds
- Guidance and Teaching Network for Video Salient Object Detection