Gunrock: A High-Performance Graph Processing Library on the GPU
arXiv:1501.05387 · doi:10.1145/2851141.2851145
Abstract
For large-scale graph analytics on the GPU, the irregularity of data access and control flow, and the complexity of programming GPUs have been two significant challenges for developing a programmable high-performance graph library. "Gunrock", our graph-processing system designed specifically for the GPU, uses a high-level, bulk-synchronous, data-centric abstraction focused on operations on a vertex or edge frontier. Gunrock achieves a balance between performance and expressiveness by coupling high performance GPU computing primitives and optimization strategies with a high-level programming model that allows programmers to quickly develop new graph primitives with small code size and minimal GPU programming knowledge. We evaluate Gunrock on five key graph primitives and show that Gunrock has on average at least an order of magnitude speedup over Boost and PowerGraph, comparable performance to the fastest GPU hardwired primitives, and better performance than any other GPU high-level graph library.
14 pages, accepted by PPoPP'16 (removed the text repetition in the previous version v5)
Cited by in corpus (19)
- The Evolution of Distributed Systems for Graph Neural Networks and their Origin in Graph Processing and Deep Learning: A Survey
- C-SAW: A Framework for Graph Sampling and Random Walk on GPUs
- A Comparative Study on Exact Triangle Counting Algorithms on the GPU
- SeGraM: A Universal Hardware Accelerator for Genomic Sequence-to-Graph and Sequence-to-Sequence Mapping
- Capstan: A Vector RDA for Sparsity
- An analysis of the graph processing landscape
- GPU Multisplit: an extended study of a parallel algorithm
- Efficient and High-quality Sparse Graph Coloring on the GPU
- Specifying and Testing GPU Workgroup Progress Models
- Bridging Control-Centric and Data-Centric Optimization
- Fast Simulation of Crowd Collision Avoidance
- Accelerating Backward Aggregation in GCN Training with Execution Path Preparing on GPUs
- DAWN: Matrix Operation-Optimized Algorithm for Shortest Paths Problem on Unweighted Graphs
- gSuite: A Flexible and Framework Independent Benchmark Suite for Graph Neural Network Inference on GPUs
- A Partition-centric Distributed Algorithm for Identifying Euler Circuits in Large Graphs
- Faster Vertex Cover Algorithms on GPUs with Component-Aware Parallel Branching
- Efficient Stepping Algorithms and Implementations for Parallel Shortest Paths
- GVEL: Fast Graph Loading in Edgelist and Compressed Sparse Row (CSR) formats
- Lifting C Semantics for Dataflow Optimization