6 citations · 6 across the 4 of their papers we have counts for
3 papers · 1 filter
Ps and Qs: Quantization-aware pruning for efficient low latency neural network inference
Benjamin Hawks, Javier Duarte, Nicholas J. Fraser +3
Efficient machine learning implementations optimized for inference in hardware have wide-ranging benefits, depending on the application, from lower inference latency to higher data…
FAT: Training Neural Networks for Reliable Inference Under Hardware Faults
Ussama Zahid, Giulio Gambardella, Nicholas J. Fraser +2
Deep neural networks (DNNs) are state-of-the-art algorithms for multiple applications, spanning from image classification to speech recognition. While providing excellent accuracy,…
Quantizing Convolutional Neural Networks for Low-Power High-Throughput Inference Engines
Sean O. Settle, Manasa Bollavaram, Paolo D'Alberto +6
Deep learning as a means to inferencing has proliferated thanks to its versatility and ability to approach or exceed human-level accuracy. These computational models have seemingly…