Benchmarking the Robustness of Instance Segmentation Models
arXiv:2109.01123 · doi:10.1109/TNNLS.2023.3310985
Abstract
This paper presents a comprehensive evaluation of instance segmentation models with respect to real-world image corruptions as well as out-of-domain image collections, e.g. images captured by a different set-up than the training dataset. The out-of-domain image evaluation shows the generalization capability of models, an essential aspect of real-world applications and an extensively studied topic of domain adaptation. These presented robustness and generalization evaluations are important when designing instance segmentation models for real-world applications and picking an off-the-shelf pretrained model to directly use for the task at hand. Specifically, this benchmark study includes state-of-the-art network architectures, network backbones, normalization layers, models trained starting from scratch versus pretrained networks, and the effect of multi-task training on robustness and generalization. Through this study, we gain several insights. For example, we find that group normalization enhances the robustness of networks across corruptions where the image contents stay the same but corruptions are added on top. On the other hand, batch normalization improves the generalization of the models across different datasets where statistics of image features change. We also find that single-stage detectors do not generalize well to larger image resolutions than their training size. On the other hand, multi-stage detectors can easily be used on images of different sizes. We hope that our comprehensive study will motivate the development of more robust and reliable instance segmentation models.
References in corpus (9)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Learning Transferable Features with Deep Adaptation Networks
- Deep Domain Confusion: Maximizing for Domain Invariance
- FCNs in the Wild: Pixel-level Adversarial and Constraint-based Adaptation
- Rethinking Pre-training and Self-training
- Convolutional Generation of Textured 3D Meshes
- Increasing the Robustness of Semantic Segmentation Models with Painting-by-Numbers
- Inst-Inpaint: Instructing to Remove Objects with Diffusion Models
- Refining 3D Human Texture Estimation from a Single Image