2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.DC2023
MCR-DL: Mix-and-Match Communication Runtime for Deep Learning
Quentin Anthony, Ammar Ahmad Awan, Jeff Rasley +5
In recent years, the training requirements of many state-of-the-art Deep Learning (DL) models have scaled beyond the compute and memory capabilities of a single processor, and nece…
cs.PF2023★ 2 cited
Performance Characterization of using Quantization for DNN Inference on Edge Devices: Extended Version
Hyunho Ahn, Tian Chen, Nawras Alnaasan +5
Quantization is a popular technique used in Deep Neural Networks (DNN) inference to reduce the size of models and improve the overall numerical performance by exploiting native har…