The MADlib Analytics Library or MAD Skills, the SQL
arXiv:1208.4165
Abstract
MADlib is a free, open source library of in-database analytic methods. It provides an evolving suite of SQL-based algorithms for machine learning, data mining and statistics that run at scale within a database engine, with no need for data import/export to other tools. The goal is for MADlib to eventually serve a role for scalable database systems that is similar to the CRAN library for R: a community repository of statistical methods, this time written with scale and parallelism in mind. In this paper we introduce the MADlib project, including the background that led to its beginnings, and the motivation for its open source nature. We provide an overview of the library's architecture and design patterns, and provide a description of various statistical methods in that context. We include performance and speedup results of a core design pattern from one of those methods over the Greenplum parallel DBMS on a modest-sized test cluster. We then report on two initial efforts at incorporating academic research into MADlib, which is one of the project's goals. MADlib is freely available at http://madlib.net, and the project is open for contributions of both new methods, and ports to additional database platforms.
VLDB2012
References in corpus (2)
Cited by in corpus (13)
- BoostClean: Automated Error Detection and Repair for Machine Learning
- LaraDB: A Minimalist Kernel for Linear and Relational Algebra Computation
- Query Processing on Tensor Computation Runtimes
- MLI: An API for Distributed Machine Learning
- Cloudy with high chance of DBMS: A 10-year prediction for Enterprise-Grade ML
- TuPAQ: An Efficient Planner for Large-scale Predictive Analytic Queries
- SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle
- Asynchronous Complex Analytics in a Distributed Dataflow Architecture
- Helix: Holistic Optimization for Accelerating Iterative Machine Learning
- Towards Linear Algebra over Normalized Data
- Are Key-Foreign Key Joins Safe to Avoid when Learning High-Capacity Classifiers?
- A Closer Look at Variance Implementations in Modern Database Systems
- Towards Expectation-Maximization by SQL in RDBMS