Gravitational tree-code on graphics processing units: implementation in CUDA
arXiv:1005.5384 · doi:10.1016/j.procs.2010.04.124
Abstract
We present a new very fast tree-code which runs on massively parallel Graphical Processing Units (GPU) with NVIDIA CUDA architecture. The tree-construction and calculation of multipole moments is carried out on the host CPU, while the force calculation which consists of tree walks and evaluation of interaction list is carried out on the GPU. In this way we achieve a sustained performance of about 100GFLOP/s and data transfer rates of about 50GB/s. It takes about a second to compute forces on a million particles with an opening angle of . The code has a convenient user interface and is freely available for use\footnote{\tt http://castle.strw.leidenuniv.nl/software/octgrav.html}.
9 pages, 8 figures. Accepted for publication at International Conference on Computational Science 2010
References in corpus (4)
- High Performance Direct Gravitational N-body Simulations on Graphics Processing Units -- II: An implementation in CUDA
- Performance Analysis of Direct N-Body Algorithms on Special-Purpose Supercomputers
- High Performance Direct Gravitational N-body Simulations on Graphics Processing Unit I: An implementation in Cg
- The Chamomile Scheme: An Optimized Algorithm for N-body simulations on Programmable Graphics Processing Units