15 citations · 16 across the 2 of their papers we have counts for
5 papers
APack: Off-Chip, Lossless Data Compression for Efficient Deep Learning Inference
Alberto Delmas Lascorz, Mostafa Mahmoud, Andreas Moshovos
Data accesses between on- and off-chip memories account for a large fraction of overall energy consumption during inference with deep learning networks. We present APack, a simple…
BitPruning: Learning Bitlengths for Aggressive and Accurate Quantization
Miloš Nikolić, Ghouthi Boukli Hacene, Ciaran Bannon +5
Neural networks have demonstrably achieved state-of-the art accuracy using low-bitlength integer quantization, yielding both execution time and energy benefits on existing hardware…
Laconic Deep Learning Computing
Sayeh Sharify, Mostafa Mahmoud, Alberto Delmas Lascorz +2
We motivate a method for transparently identifying ineffectual computations in unmodified Deep Learning models and without affecting accuracy. Specifically, we show that if we deco…
Loom: Exploiting Weight and Activation Precisions to Accelerate Convolutional Neural Networks
Sayeh Sharify, Alberto Delmas Lascorz, Kevin Siu +2
Loom (LM), a hardware inference accelerator for Convolutional Neural Networks (CNNs) is presented. In LM every bit of data precision that can be saved translates to proportional pe…
Cnvlutin2: Ineffectual-Activation-and-Weight-Free Deep Neural Network Computing
Patrick Judd, Alberto Delmas, Sayeh Sharify +1
We discuss several modifications and extensions over the previous proposed Cnvlutin (CNV) accelerator for convolutional and fully-connected layers of Deep Learning Network. We firs…