6 citations · 10 across the 5 of their papers we have counts for
3 papers · 1 filter
The Case for Co-Designing Model Architectures with Hardware
Quentin Anthony, Jacob Hatef, Deepak Narayanan +6
While GPUs are responsible for training the vast majority of state-of-the-art deep learning models, the implications of their architecture are often overlooked when designing new d…
MCR-DL: Mix-and-Match Communication Runtime for Deep Learning
Quentin Anthony, Ammar Ahmad Awan, Jeff Rasley +5
In recent years, the training requirements of many state-of-the-art Deep Learning (DL) models have scaled beyond the compute and memory capabilities of a single processor, and nece…
System-level Scalable Checkpoint-Restart for Petascale Computing
Jiajun Cao, Kapil Arya, Rohan Garg +5
Fault tolerance for the upcoming exascale generation has long been an area of active research. One of the components of a fault tolerance strategy is checkpointing. Petascale-level…