40 citations · 56 across the 2 of their papers we have counts for
2 papers
cs.DC2016★ 40 cited
PANDA: Extreme Scale Parallel K-Nearest Neighbor on Distributed Architectures
Md. Mostofa Ali Patwary, Nadathur Rajagopalan Satish, Narayanan Sundaram +8
Computing -Nearest Neighbors (KNN) is one of the core kernels used in many machine learning, data mining and scientific computing applications. Although kd-tree based $O(\log n)…
cs.DC2016★ 16 cited
Matrix Factorization at Scale: a Comparison of Scientific Data Analytics in Spark and C+MPI Using Three Case Studies
Alex Gittens, Aditya Devarakonda, Evan Racah +14
We explore the trade-offs of performing linear algebra using Apache Spark, compared to traditional C and MPI implementations on HPC platforms. Spark is designed for data analytics…