SegDiff: Image Segmentation with Diffusion Probabilistic Models
arXiv:2112.00390
Abstract
Diffusion Probabilistic Methods are employed for state-of-the-art image generation. In this work, we present a method for extending such models for performing image segmentation. The method learns end-to-end, without relying on a pre-trained backbone. The information in the input image and in the current estimation of the segmentation map is merged by summing the output of two encoders. Additional encoding layers and a decoder are then used to iteratively refine the segmentation map, using a diffusion model. Since the diffusion model is probabilistic, it is applied multiple times, and the results are merged into a final segmentation map. The new method produces state-of-the-art results on the Cityscapes validation set, the Vaihingen building segmentation benchmark, and the MoNuSeg dataset.
References in corpus (14)
- Diffusion Models Beat GANs on Image Synthesis
- SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers
- Semantic Segmentation using Adversarial Networks
- Cascaded Diffusion Models for High Fidelity Image Generation
- Improved Denoising Diffusion Probabilistic Models
- Variational Diffusion Models
- DiffWave: A Versatile Diffusion Model for Audio Synthesis
- Structured Denoising Diffusion Models in Discrete State-Spaces
- Adversarial Examples for Semantic Image Segmentation
- Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech
- A Variational Perspective on Diffusion-Based Generative Models and Score Matching
- Argmax Flows and Multinomial Diffusion: Learning Categorical Distributions
- SRDiff: Single Image Super-Resolution with Diffusion Probabilistic Models
- End to End Trainable Active Contours via Differentiable Rendering