PyCUDA and PyOpenCL: A Scripting-Based Approach to GPU Run-Time Code Generation
arXiv:0911.3456 · doi:10.1016/j.parco.2011.09.001
Abstract
High-performance computing has recently seen a surge of interest in heterogeneous systems, with an emphasis on modern Graphics Processing Units (GPUs). These devices offer tremendous potential for performance and efficiency in important large-scale applications of computational science. However, exploiting this potential can be challenging, as one must adapt to the specialized and rapidly evolving computing environment currently exhibited by GPUs. One way of addressing this challenge is to embrace better techniques and develop tools tailored to their needs. This article presents one simple technique, GPU run-time code generation (RTCG), along with PyCUDA and PyOpenCL, two open-source toolkits that support this technique. In introducing PyCUDA and PyOpenCL, this article proposes the combination of a dynamic, high-level scripting language with the massive performance of a GPU as a compelling two-tiered computing platform, potentially offering significant performance and productivity advantages over conventional single-tier, static systems. The concept of RTCG is simple and easily implemented using existing, robust infrastructure. Nonetheless it is powerful enough to support (and encourage) the creation of custom application-specific tools by its users. The premise of the paper is illustrated by a wide range of examples where the technique has been applied with considerable success.
Submitted to Parallel Computing, Elsevier
References in corpus (1)
Cited by in corpus (102)
- PyFR: An Open Source Framework for Solving Advection-Diffusion Type Problems on Streaming Architectures using the Flux Reconstruction Approach
- Chaotic lensing around boson stars and Kerr black holes with scalar hair
- HELIOS: An Open-source, GPU-accelerated Radiative Transfer Code For Self-consistent Exoplanetary Atmospheres
- Programming Languages for Scientific Computing
- Self-luminous and irradiated exoplanetary atmospheres explored with HELIOS
- Sailfish: a flexible multi-GPU implementation of the lattice Boltzmann method
- Constraining axion inflation with gravitational waves from preheating
- Maximum a posteriori CMB lensing reconstruction
- Constraining axion inflation with gravitational waves across 29 decades in frequency
- Hydrodynamics of Suspensions of Passive and Active Rigid Particles: A Rigid Multiblob Approach
- All-particle cosmic ray energy spectrum measured by the HAWC experiment from 10 to 500 TeV
- SMUTHI: A python package for the simulation of light scattering by multiple particles near or between planar interfaces
- Image Subtraction in Fourier Space
- PySPH: a Python-based framework for smoothed particle hydrodynamics
- A biomolecular electrostatics solver using Python, GPUs and boundary elements that can handle solvent-filled cavities and Stern layers
- Loo.py: transformation-based code generation for GPUs and CPUs
- Brownian Dynamics of Confined Suspensions of Active Microrollers
- The Detectability of Rocky Planet Surface and Atmosphere Composition with JWST: The Case of LHS 3844b
- PeriPy -- A High Performance OpenCL Peridynamics Package
- THOR 2.0: Major Improvements to the Open-Source General Circulation Model
- QInfer: Statistical inference software for quantum applications
- Evaluation of OpenAI Codex for HPC Parallel Programming Models Kernel Generation
- Bayesian Neural Networks for Genetic Association Studies of Complex Disease
- Theano-based Large-Scale Visual Recognition with Multiple GPUs
- Neutron star mass estimates from gamma-ray eclipses in spider millisecond pulsar binaries
- Constraining early dark energy with gravitational waves before recombination
- Biobeam - Rigorous wave-optical simulations of light-sheet microscopy
- PyCOOL - a Cosmological Object-Oriented Lattice code written in Python
- Measuring the mass of the black widow PSR J1555-2908
- Deep Learning for Ontology Reasoning
- Transiting Planets near the Snow Line from Kepler. I. Catalog
- Chainer: A Deep Learning Framework for Accelerating the Research Cycle
- Design and implementation of a multi-octave-band audio camera for realtime diagnosis
- PyFstat: a Python package for continuous gravitational-wave data analysis
- Montblanc: GPU accelerated Radio Interferometer Measurement Equations in support of Bayesian Inference for Radio Observations
- Accelerated Matrix Element Method with Parallel Computing
- A Mesoscale Perspective on the Tolman Length
- GPU Computing with Python: Performance, Energy Efficiency and Usability
- Evaluating performance and portability of high-level programming models: Julia, Python/Numba, and Kokkos on exascale nodes
- First measurement of the -violating phase in decays
- Gauge preheating with full general relativity
- Spontaneous Symmetry Breaking for Extreme Vorticity and Strain in the 3D Navier-Stokes Equations
- Faster search for long gravitational-wave transients: GPU implementation of the transient F-statistic
- USID and Pycroscopy -- Open frameworks for storing and analyzing spectroscopic and imaging data
- Nonperturbative structure in coupled axion sectors and implications for direct detection
- Dirac open quantum system dynamics: formulations and simulations
- Resolving the blazar CGRaBS J0809+5341 in the presence of telescope systematics
- High-Performance Multi-Mode Ptychography Reconstruction on Distributed GPUs
- PyMatting: A Python Library for Alpha Matting
- Structure and Isotropy of Lattice Pressure Tensors for Multi-range Potentials
- Performance Evaluation of Python Parallel Programming Models: Charm4Py and mpi4py
- Large-scale comparative visualisation of sets of multidimensional data
- Normal and anomalous random walks of 2-d solitons
- Parameter space metric for 3.5 post-Newtonian gravitational-waves from compact binary inspirals
- Optimal Experimental Design for Uncertain Systems Based on Coupled Differential Equations
- High Level Programming for Heterogeneous Architectures
- Using SIMD and SIMT vectorization to evaluate sparse chemical kinetic Jacobian matrices and thermochemical source terms
- Validity of Gross-Pitaevskii solutions of harmonically confined BEC gases in reduced dimensions
- Efficiency and accuracy of GPU-parallelized Fourier spectral methods for solving phase-field models
- Measuring Star-Formation Histories, Distances, and Metallicities with Pixel Color-Magnitude Diagrams I: Model Definition and Mock Tests
- Study of the decay with an amplitude analysis of decays
- Measuring Star Formation Histories, Distances, and Metallicities with Pixel Color-Magnitude Diagrams II: Applications to Nearby Elliptical Galaxies
- Scalable Methods for Computing Sharp Extreme Event Probabilities in Infinite-Dimensional Stochastic Systems
- Rapid Development of Interferometric Software Using MIRIAD and Python
- PyWolf: A PyOpenCL implementation for simulating the propagation of partially coherent light
- Chaotic Diffusion of Dissipative Solitons: From Anti-Persistent Random Walk to Hidden Markov Models
- Non-linear fitting with joint spatial regularization in Arterial Spin Labeling
- Contract-Based General-Purpose GPU Programming
- Fluid dynamics in porous media with Sailfish
- CUDA-Self-Organizing feature map based visual sentiment analysis of bank customer complaints for Analytical CRM
- Genetic studies through the lens of gene networks
- Robust Statistics for Image Deconvolution
- Formal Definition and Implementation of Reproducibility Tenets for Computational Workflows
- OpenCL Actors - Adding Data Parallelism to Actor-based Programming with CAF
- Casimir Forces via Worldline Numerics: Method Improvements and Potential Engineering Applications
- Comparing Llama-2 and GPT-3 LLMs for HPC kernels generation
- Bringing Together Dynamic Geometry Software and the Graphics Processing Unit
- DistStat.jl: Towards Unified Programming for High-Performance Statistical Computing Environments in Julia
- The Monte Carlo Computational Summit -- October 25 & 26, 2023 -- Notre Dame, Indiana, USA
- Reproducible Validation and Replication Studies in Nanoscale Physics
- Fast Hamiltonian Monte Carlo Using GPU Computing
- cf4ocl: a C framework for OpenCL
- Python Non-Uniform Fast Fourier Transform (PyNUFFT): multi-dimensional non-Cartesian image reconstruction package for heterogeneous platforms and applications to MRI
- Scaling and Acceleration of Three-dimensional Structure Determination for Single-Particle Imaging Experiments with SpiniFEL
- Convolutions of Totally Positive Distributions with applications to Kernel Density Estimation
- Towards Understanding Residual and Dilated Dense Neural Networks via Convolutional Sparse Coding
- Real-Time Rendering of Arbitrary Surface Geometries using Learnt Transfer
- Hiperwalk: Simulation of Quantum Walks with Heterogeneous High-Performance Computing
- Position Paper: Towards Transparent Machine Learning
- OpenCL-accelerated object classification in video streams using Spatial Pooler of Hierarchical Temporal Memory
- High-Order Discontinuous Galerkin Methods by GPU Metaprogramming
- Composition-Aware Spectroscopic Tomography
- Real-World Oceanographic Simulations on the GPU using a Two-Dimensional Finite-Volume Scheme
- Lya2pcf: an efficient pipeline to estimate two- and three-point correlation functions of the Lyman- forest
- The Two-Dimensional Swept Rule Applied on Heterogeneous Architectures
- On new data sources for the production of official statistics
- Real time computer generation of three-dimensional point cloud holograms through GPU implementation of compressed sensing Gerchberg-Saxton algorithm
- Massively parallel read mapping on GPUs with PEANUT
- An ensemble solver for segregated cardiovascular FSI
- Fireflies: New software for interactively exploring dynamical systems using GPU computing
- Energy-Efficient p-Bit-Based Fully-Connected Quantum-Inspired Simulated Annealer with Dual BRAM Architecture
- Heterogeneous Computing on Mixed Unstructured Grids with PyFR