2 papers
cs.LG2021
Ps and Qs: Quantization-aware pruning for efficient low latency neural network inference
Benjamin Hawks, Javier Duarte, Nicholas J. Fraser +3
Efficient machine learning implementations optimized for inference in hardware have wide-ranging benefits, depending on the application, from lower inference latency to higher data…
cs.CV2018
FINN-L: Library Extensions and Design Trade-off Analysis for Variable Precision LSTM Networks on FPGAs
Vladimir Rybalkin, Alessandro Pappalardo, Muhammad Mohsin Ghaffar +3
It is well known that many types of artificial neural networks, including recurrent networks, can achieve a high classification accuracy even with low-precision weights and activat…