Graphulo: Linear Algebra Graph Kernels for NoSQL Databases
arXiv:1508.07372 · doi:10.1109/IPDPSW.2015.19
Abstract
Big data and the Internet of Things era continue to challenge computational systems. Several technology solutions such as NoSQL databases have been developed to deal with this challenge. In order to generate meaningful results from large datasets, analysts often use a graph representation which provides an intuitive way to work with the data. Graph vertices can represent users and events, and edges can represent the relationship between vertices. Graph algorithms are used to extract meaningful information from these very large graphs. At MIT, the Graphulo initiative is an effort to perform graph algorithms directly in NoSQL databases such as Apache Accumulo or SciDB, which have an inherently sparse data storage scheme. Sparse matrix operations have a history of efficient implementations and the Graph Basic Linear Algebra Subprogram (GraphBLAS) community has developed a set of key kernels that can be used to develop efficient linear algebra operations. However, in order to use the GraphBLAS kernels, it is important that common graph algorithms be recast using the linear algebra building blocks. In this article, we look at common classes of graph algorithms and recast them into linear algebra operations using the GraphBLAS building blocks.
10 pages
References in corpus (2)
Cited by in corpus (15)
- Static Graph Challenge: Subgraph Isomorphism
- LaraDB: A Minimalist Kernel for Linear and Relational Algebra Computation
- Graphulo Implementation of Server-Side Sparse Matrix Multiply in the Accumulo Database
- D4M: Bringing Associative Arrays to Database Engines
- Polystore Mathematics of Relational Algebra
- From NoSQL Accumulo to NewSQL Graphulo: Design and Utility of Graph Algorithms inside a BigTable Database
- Streaming 1.9 Billion Hypersparse Network Updates per Second with D4M
- In-Storage Embedded Accelerator for Sparse Pattern Processing
- Distributed Triangle Counting in the Graphulo Matrix Math Library
- Mathematics of Digital Hyperspace
- Benchmarking the Graphulo Processing Framework
- TabulaROSA: Tabular Operating System Architecture for Massively Parallel Heterogeneous Compute Engines
- D4M 3.0: Extended Database and Language Capabilities
- Python Implementation of the Dynamic Distributed Dimensional Data Model
- Large Enforced Sparse Non-Negative Matrix Factorization