PRAND: GPU accelerated parallel random number generation library: Using most reliable algorithms and applying parallelism of modern GPUs and CPUs
arXiv:1307.5869 · doi:10.1016/j.cpc.2014.01.007
Abstract
The library PRAND for pseudorandom number generation for modern CPUs and GPUs is presented. It contains both single-threaded and multi-threaded realizations of a number of modern and most reliable generators recently proposed and studied in [1,2,3,4,5] and the efficient SIMD realizations proposed in [6]. One of the useful features for using PRAND in parallel simulations is the ability to initialize up to independent streams. Using massive parallelism of modern GPUs and SIMD parallelism of modern CPUs substantially improves performance of the generators.
29 pages, 1 figure, 7 tables
References in corpus (6)
- Random numbers for large scale distributed Monte Carlo simulations
- Random number generators for massively parallel simulations on GPU
- Pseudo-random number generators for Monte Carlo simulations on Graphics Processing Units
- RNGSSELIB: Program library for random number generation, SSE2 realization
- RNGSSELIB: Program library for random number generation. More generators, parallel streams of random numbers and Fortran compatibility
- Applying dissipative dynamical systems to pseudorandom number generation: Equidistribution property and statistical independence of bits at distances up to logarithm of mesh size
Cited by in corpus (9)
- GPU accelerated population annealing algorithm
- Competing nematic interactions in a generalized XY model in two and three dimensions
- Massively parallel multicanonical simulations
- Joint effect of advection, diffusion, and capillary attraction on the spatial structure of particle depositions from evaporating droplets
- Highly optimized simulations on single- and multi-GPU systems of 3D Ising spin glass
- RNGSSELIB: Program library for random number generation. More generators, parallel streams of random numbers and Fortran compatibility
- High-Performance Hybrid Algorithm for Minimum Sum-of-Squares Clustering of Infinitely Tall Data
- Algorithm for the replica redistribution in the implementation of parallel annealing method on the hybrid supercomputer architecture
- Comparison of the microcanonical population annealing algorithm with the Wang-Landau algorithm