Graphulo Implementation of Server-Side Sparse Matrix Multiply in the Accumulo Database
arXiv:1507.01066 · doi:10.1109/HPEC.2015.7322448
Abstract
The Apache Accumulo database excels at distributed storage and indexing and is ideally suited for storing graph data. Many big data analytics compute on graph data and persist their results back to the database. These graph calculations are often best performed inside the database server. The GraphBLAS standard provides a compact and efficient basis for a wide range of graph applications through a small number of sparse matrix operations. In this article, we implement GraphBLAS sparse matrix multiplication server-side by leveraging Accumulo's native, high-performance iterators. We compare the mathematics and performance of inner and outer product implementations, and show how an outer product implementation achieves optimal performance near Accumulo's peak write rate. We offer our work as a core component to the Graphulo library that will deliver matrix math primitives for graph analytics within Accumulo.
To be presented at IEEE HPEC 2015: http://www.ieee-hpec.org/
References in corpus (3)
Cited by in corpus (20)
- Mathematical Foundations of the GraphBLAS
- The BigDAWG Polystore System and Architecture
- A Systematic Survey of General Sparse Matrix-Matrix Multiplication
- LaraDB: A Minimalist Kernel for Linear and Relational Algebra Computation
- Polystore Mathematics of Relational Algebra
- From NoSQL Accumulo to NewSQL Graphulo: Design and Utility of Graph Algorithms inside a BigTable Database
- Streaming 1.9 Billion Hypersparse Network Updates per Second with D4M
- In-Storage Embedded Accelerator for Sparse Pattern Processing
- BigSparse: High-performance external graph analytics
- Distributed Triangle Counting in the Graphulo Matrix Math Library
- A Billion Updates per Second Using 30,000 Hierarchical In-Memory D4M Databases
- Mathematics of Digital Hyperspace
- Constructing Adjacency Arrays from Incidence Arrays
- Benchmarking the Graphulo Processing Framework
- TabulaROSA: Tabular Operating System Architecture for Massively Parallel Heterogeneous Compute Engines
- SoK: Cryptographically Protected Database Search
- D4M 3.0: Extended Database and Language Capabilities
- The BigDAWG Architecture
- High-throughput Ingest of Provenance Records into Accumulo
- SAGE: A Storage-Based Approach for Scalable and Efficient Sparse Generalized Matrix-Matrix Multiplication