438 citations · 508 across the 12 of their papers we have counts for
5 papers · 1 filter
Alchemist: An Apache Spark <=> MPI Interface
Alex Gittens, Kai Rothauge, Shusen Wang +6
The Apache Spark framework for distributed computation is popular in the data analytics community due to its ease of use, but its MapReduce-style programming model can incur signif…
Accelerating Large-Scale Data Analysis by Offloading to High-Performance Computing Libraries using Alchemist
Alex Gittens, Kai Rothauge, Shusen Wang +6
Apache Spark is a popular system aimed at the analysis of large data sets, but recent studies have shown that certain computations---in particular, many linear algebra computations…
Cataloging the Visible Universe through Bayesian Inference at Petascale
Jeffrey Regier, Kiran Pamnany, Keno Fischer +9
Astronomical catalogs derived from wide-field imaging surveys are an important tool for understanding the Universe. We construct an astronomical catalog from 55 TB of imaging data…
Scaling GRPC Tensorflow on 512 nodes of Cori Supercomputer
Amrita Mathuriya, Thorsten Kurth, Vivek Rane +5
We explore scaling of the standard distributed Tensorflow with GRPC primitives on up to 512 Intel Xeon Phi (KNL) nodes of Cori supercomputer with synchronous stochastic gradient de…
An Assessment of Data Transfer Performance for Large-Scale Climate Data Analysis and Recommendations for the Data Infrastructure for CMIP6
Eli Dart, Michael F. Wehner, Prabhat
We document the data transfer workflow, data transfer performance, and other aspects of staging approximately 56 terabytes of climate model output data from the distributed Coupled…