Pushing the limits for medical image reconstruction on recent standard multicore processors
arXiv:1104.5243 · doi:10.1177/1094342012442424
Abstract
Volume reconstruction by backprojection is the computational bottleneck in many interventional clinical computed tomography (CT) applications. Today vendors in this field replace special purpose hardware accelerators by standard hardware like multicore chips and GPGPUs. Medical imaging algorithms are on the verge of employing High Performance Computing (HPC) technology, and are therefore an interesting new candidate for optimization. This paper presents low-level optimizations for the backprojection algorithm, guided by a thorough performance analysis on four generations of Intel multicore processors (Harpertown, Westmere, Westmere EX, and Sandy Bridge). We choose the RabbitCT benchmark, a standardized testcase well supported in industry, to ensure transparent and comparable results. Our aim is to provide not only the fastest possible implementation but also compare to performance models and hardware counter data in order to fully understand the results. We separate the influence of algorithmic optimizations, parallelization, SIMD vectorization, and microarchitectural issues and pinpoint problems with current SIMD instruction set extensions on standard CPUs (SSE, AVX). The use of assembly language is mandatory for best performance. Finally we compare our results to the best GPGPU implementations available for this open competition benchmark.
13 pages, 9 figures. Revised and extended version
References in corpus (3)
Cited by in corpus (10)
- Exploring performance and power properties of modern multicore chips via simple machine models
- Chip-level and multi-node analysis of energy-optimized lattice-Boltzmann CFD simulations
- Comparing the Performance of Different x86 SIMD Instruction Sets for a Medical Imaging Application on Modern Multi- and Manycore Chips
- Petascale XCT: 3D Image Reconstruction with Hierarchical Communications on Multi-GPU Nodes
- Best practices for HPM-assisted performance engineering on modern multicore processors
- iFDK: A Scalable Framework for Instant High-resolution Image Reconstruction
- Performance Engineering for a Medical Imaging Application on the Intel Xeon Phi Accelerator
- Performance analysis of the Kahan-enhanced scalar product on current multi- and manycore processors
- Performance Portable Back-projection Algorithms on CPUs: Agnostic Data Locality and Vectorization Optimizations
- Deep Learning Accelerated Light Source Experiments