16 citations · 61 across the 5 of their papers we have counts for
5 papers
Large-Batch Training for LSTM and Beyond
Yang You, Jonathan Hseu, Chris Ying +3
Large-batch training approaches have enabled researchers to utilize large-scale distributed processing and greatly accelerate deep-neural net (DNN) training. For example, by scalin…
Parallelepipeds obtaining HBL lower bounds
James Demmel, Alex Rusciano
This work studies the application of the discrete Holder-Brascamp-Lieb (HBL) inequalities to the design of communication optimal algorithms. In particular, it describes optimal til…
Matrix Factorization at Scale: a Comparison of Scientific Data Analytics in Spark and C+MPI Using Three Case Studies
Alex Gittens, Aditya Devarakonda, Evan Racah +14
We explore the trade-offs of performing linear algebra using Apache Spark, compared to traditional C and MPI implementations on HPC platforms. Spark is designed for data analytics…
Strong Scaling of Matrix Multiplication Algorithms and Memory-Independent Communication Lower Bounds
Grey Ballard, James Demmel, Olga Holtz +2
A parallel algorithm has perfect strong scaling if its running time on P processors is linear in 1/P, including all communication costs. Distributed-memory parallel algorithms for…
Communication-Optimal Parallel Algorithm for Strassen's Matrix Multiplication
Grey Ballard, James Demmel, Olga Holtz +2
Parallel matrix multiplication is one of the most studied fundamental problems in distributed and high performance computing. We obtain a new parallel algorithm that is based on St…