D4M: Bringing Associative Arrays to Database Engines
arXiv:1508.07371 · doi:10.1109/HPEC.2015.7322472
Abstract
The ability to collect and analyze large amounts of data is a growing problem within the scientific community. The growing gap between data and users calls for innovative tools that address the challenges faced by big data volume, velocity and variety. Numerous tools exist that allow users to store, query and index these massive quantities of data. Each storage or database engine comes with the promise of dealing with complex data. Scientists and engineers who wish to use these systems often quickly find that there is no single technology that offers a panacea to the complexity of information. When using multiple technologies, however, there is significant trouble in designing the movement of information between storage and database engines to support an end-to-end application along with a steep learning curve associated with learning the nuances of each underlying technology. In this article, we present the Dynamic Distributed Dimensional Data Model (D4M) as a potential tool to unify database and storage engine operations. Previous articles on D4M have showcased the ability of D4M to interact with the popular NoSQL Accumulo database. Recently however, D4M now operates on a variety of backend storage or database engines while providing a federated look to the end user through the use of associative arrays. In order to showcase how new databases may be supported by D4M, we describe the process of building the D4M-SciDB connector and present performance of this connection.
References in corpus (2)
Cited by in corpus (15)
- Interactive Supercomputing on 40,000 Cores for Machine Learning and Data Analysis
- The BigDAWG Polystore System and Architecture
- Measuring the Impact of Spectre and Meltdown
- Polystore Mathematics of Relational Algebra
- LLMapReduce: Multi-Level Map-Reduce for High Performance Data Analysis
- Benchmarking Data Analysis and Machine Learning Applications on the Intel KNL Many-Core Processor
- Node-Based Job Scheduling for Large Scale Simulations of Short Running Jobs
- Maneuver Identification Challenge
- Lessons Learned from a Decade of Providing Interactive, On-Demand High Performance Computing to Scientists and Engineers
- Mathematics of Digital Hyperspace
- Best of Both Worlds: High Performance Interactive and Batch Launching
- Securing HPC using Federated Authentication
- Python Implementation of the Dynamic Distributed Dimensional Data Model
- HPC with Enhanced User Separation
- Linear Systems over Join-Blank Algebras