961 citations · 980 across the 3 of their papers we have counts for
3 papers
cs.LG2015★ 961 cited
MLlib: Machine Learning in Apache Spark
Xiangrui Meng, Joseph Bradley, Burak Yavuz +13
Apache Spark is a popular open-source platform for large-scale data processing that is well-suited for iterative machine learning tasks. In this paper we present MLlib, Spark's ope…
cs.DB2012★ 19 cited
Shark: SQL and Rich Analytics at Scale
Reynold Xin, Josh Rosen, Matei Zaharia +3
Shark is a new data analysis system that marries query processing with complex analytics on large clusters. It leverages a novel distributed memory abstraction to provide a unified…
cs.DB2012
The End of an Architectural Era for Analytical Databases
Reynold S. Xin
Traditional enterprise warehouse solutions center around an analytical database system that is monolithic and inflexible: data needs to be extracted, transformed, and loaded into t…