A sparse octree gravitational N-body code that runs entirely on the GPU processor
arXiv:1106.1900 · doi:10.1016/j.jcp.2011.12.024
Abstract
We present parallel algorithms for constructing and traversing sparse octrees on graphics processing units (GPUs). The algorithms are based on parallel-scan and sort methods. To test the performance and feasibility, we implemented them in CUDA in the form of a gravitational tree-code which completely runs on the GPU.(The code is publicly available at: http://castle.strw.leidenuniv.nl/software.html) The tree construction and traverse algorithms are portable to many-core devices which have support for CUDA or OpenCL programming languages. The gravitational tree-code outperforms tuned CPU code during the tree-construction and shows a performance improvement of more than a factor 20 overall, resulting in a processing rate of more than 2.8 million particles per second.
Accepted version. Published in Journal of Computational Physics. 35 pages, 12 figures, single column
References in corpus (6)
- High Performance Direct Gravitational N-body Simulations on Graphics Processing Units -- II: An implementation in CUDA
- A multiphysics and multiscale software environment for modeling astrophysical systems
- Performance Analysis of Direct N-Body Algorithms on Special-Purpose Supercomputers
- High Performance Direct Gravitational N-body Simulations on Graphics Processing Unit I: An implementation in Cg
- Gravitational tree-code on graphics processing units: implementation in CUDA
- The Chamomile Scheme: An Optimized Algorithm for N-body simulations on Programmable Graphics Processing Units
Cited by in corpus (53)
- Multi-physics simulations using a hierarchical interchangeable software interface
- The Astrophysical Multipurpose Software Environment
- Implementation and performance of FDPS: A Framework Developing Parallel Particle Simulation Codes
- The origin of interstellar asteroidal objects like 1I/2017 U1 'Oumuamua
- 24.77 Pflops on a Gravitational Tree-Code to Simulate the Milky Way Galaxy with 18600 GPUs
- The dynamics of stellar disks in live dark-matter halos
- Multiple phase spirals suggest multiple origins in Gaia DR3
- Phantom-GRAPE: numerical software library to accelerate collisionless -body simulation with SIMD instruction set on x86 architecture
- A smooth particle hydrodynamics code to model collisions between solid, self-gravitating objects
- Modeling the Milky Way as a Dry Galaxy
- 2HOT: An Improved Parallel Hashed Oct-Tree N-Body Algorithm for Cosmological Simulation
- Evolution of star clusters in a cosmological tidal field
- Trimodal structure of Hercules stream explained by originating from bar resonances
- The Effect of Many Minor Mergers on the Size Growth of Compact Quiescent Galaxies
- Swarm-NG: a CUDA Library for Parallel n-body Integrations with focus on Simulations of Planetary Systems
- Accelerated FDPS --- Algorithms to Use Accelerators with FDPS
- Snails Across Scales: Local and Global Phase-Mixing Structures as Probes of the Past and Future Milky Way
- GOTHIC: Gravitational oct-tree code accelerated by hierarchical time step controlling
- Dynamical evolution of massive black holes in galactic-scale N-body simulations - introducing the regularized tree code "rVINE"
- A GPU-Accelerated Fast Summation Method Based on Barycentric Lagrange Interpolation and Dual Tree Traversal
- Simulations of the tidal interaction and mass transfer of a star in an eccentric orbit around an intermediate-mass black hole: the case of HLX-1
- Radial phase spirals in the Solar neighbourhood
- Ripples spreading across the Galactic disc. Interplay of direct and indirect effects of the Sagittarius dwarf impact
- Growing Local arm inferred by the breathing motion
- Weighing the Galactic disk using phase-space spirals IV. Tests on a three-dimensional galaxy simulation
- Dynamics of supermassive black hole triples in the ROMULUS25 cosmological simulation
- Impact of bar resonances in the velocity-space distribution of the solar neighbourhood stars in a self-consistent -body Galactic disc simulation
- Cornerstone: Octree Construction Algorithms for Scalable Particle Simulations
- GPU parallel simulation algorithm of Brownian particles with excluded volume using Delaunay triangulations
- A pilgrimage to gravity on GPUs
- Performance analysis of parallel gravitational -body codes on large GPU cluster
- GPU-accelerated simulation of colloidal suspensions with direct hydrodynamic interactions
- Astrophysical Particle Simulations on Heterogeneous CPU-GPU Systems
- Multi-scale and multi-domain computational astrophysics
- DeepGalaxy: Deducing the Properties of Galaxy Mergers from Images Using Deep Neural Networks
- Terrestrial planet formation during giant planet formation and giant planet migration I: The first 5 million years
- An FMM Based on Dual Tree Traversal for Many-core Architectures
- Communication Complexity of the Fast Multipole Method and its Algebraic Variants
- The effect of kick velocities on the spatial distribution of millisecond pulsars and implications for the Galactic center excess
- Parallel Algorithms for Constructing Data Structures for Fast Multipole Methods
- Delayed phase mixing in the self-gravitating Galactic disc
- Acceleration of the tree method with SIMD instruction set
- Stochastic Barnes-Hut Approximation for Fast Summation on the GPU
- Communication Reducing Algorithms for Distributed Hierarchical N-Body Problems with Boundary Distributions
- Resolving local and global kinematic signatures of satellite mergers with billion particle simulations
- The dark matter wake of a galactic bar revealed by multichannel Singular Spectral Analysis
- GPU-Enabled Particle-Particle Particle-Tree Scheme for Simulating Dense Stellar Cluster System
- Gravitational octree code performance evaluation on Volta GPU
- Cosmological Calculations on the GPU
- MatRox: Modular approach for improving data locality in Hierarchical (Mat)rix App(Rox)imation
- OpenGadget3 GPU solver tests
- SCF-FDPS: A Fast -body Code for Simulating Disk-halo Systems
- A Non-linear GPU Thread Map for Triangular Domains