3DQ: Compact Quantized Neural Networks for Volumetric Whole Brain Segmentation
arXiv:1904.03110 · doi:10.1007/978-3-030-32248-9_49
Abstract
Model architectures have been dramatically increasing in size, improving performance at the cost of resource requirements. In this paper we propose 3DQ, a ternary quantization method, applied for the first time to 3D Fully Convolutional Neural Networks (F-CNNs), enabling 16x model compression while maintaining performance on par with full precision models. We extensively evaluate 3DQ on two datasets for the challenging task of whole brain segmentation. Additionally, we showcase our method's ability to generalize on two common 3D architectures, namely 3D U-Net and V-Net. Outperforming a variety of baselines, the proposed method is capable of compressing large 3D models to a few MBytes, alleviating the storage needs in space critical applications.
Accepted to MICCAI 2019
References in corpus (5)
- Distilling the Knowledge in a Neural Network
- A Survey of Model Compression and Acceleration for Deep Neural Networks
- Trained Ternary Quantization
- A large annotated medical image dataset for the development and evaluation of segmentation algorithms
- TernaryNet: Faster Deep Model Inference without GPUs for Medical 3D Segmentation using Sparse and Binary Convolutions