Video-SwinUNet: Spatio-temporal Deep Learning Framework for VFSS Instance Segmentation
arXiv:2302.11325 · doi:10.1109/ICIP49359.2023
Abstract
This paper presents a deep learning framework for medical video segmentation. Convolution neural network (CNN) and transformer-based methods have achieved great milestones in medical image segmentation tasks due to their incredible semantic feature encoding and global information comprehension abilities. However, most existing approaches ignore a salient aspect of medical video data - the temporal dimension. Our proposed framework explicitly extracts features from neighbouring frames across the temporal dimension and incorporates them with a temporal feature blender, which then tokenises the high-level spatio-temporal feature to form a strong global feature encoded via a Swin Transformer. The final segmentation results are produced via a UNet-like encoder-decoder architecture. Our model outperforms other approaches by a significant margin and improves the segmentation benchmarks on the VFSS2022 dataset, achieving a dice coefficient of 0.8986 and 0.8186 for the two datasets tested. Our studies also show the efficacy of the temporal feature blending scheme and cross-dataset transferability of learned capabilities. Code and models are fully available at https://github.com/SimonZeng7108/Video-SwinUNet.
References in corpus (1)
Cited by in corpus (9)
- Experimental comparison of single-pixel imaging algorithms
- A review of advancements in low-light image enhancement using deep learning
- 2D Image head pose estimation via latent space regression under occlusion settings
- Topological Persistence Guided Knowledge Distillation for Wearable Sensor Data
- BGrowth: an efficient approach for the segmentation of vertebral compression fractures in magnetic resonance imaging
- JotlasNet: Joint Tensor Low-Rank and Attention-based Sparse Unrolling Network for Accelerating Dynamic MRI
- ActiveFreq: Integrating Active Learning and Frequency Domain Analysis for Interactive Segmentation
- DMSORT: An efficient parallel maritime multi-object tracking architecture for unmanned vessel platforms
- Two-Stream Thermal Imaging Fusion for Enhanced Time of Birth Detection in Neonatal Care