2 citations · 2 across the 4 of their papers we have counts for
5 papers
SIRA: Scaled-Integer Range Analysis for Optimizing FPGA Dataflow Neural Network Accelerators
Yaman Umuroglu, Christoph Berganski, Felix Jentzsch +8
While neural network quantization effectively reduces the cost of matrix multiplications, aggressive quantization can expose non-matrix-multiply operations as significant performan…
FINN-GL: Generalized Mixed-Precision Extensions for FPGA-Accelerated LSTMs
Shashwat Khandelwal, Jakoba Petri-Koenig, Thomas B. Preußer +2
Recurrent neural networks (RNNs), particularly LSTMs, are effective for time-series tasks like sentiment analysis and short-term stock prediction. However, their computational comp…
A2Q+: Improving Accumulator-Aware Weight Quantization
Ian Colbert, Alessandro Pappalardo, Jakoba Petri-Koenig +1
Quantization techniques commonly reduce the inference costs of neural networks by restricting the precision of weights and activations. Recent studies show that also reducing the p…
A2Q: Accumulator-Aware Quantization with Guaranteed Overflow Avoidance
Ian Colbert, Alessandro Pappalardo, Jakoba Petri-Koenig
We present accumulator-aware quantization (A2Q), a novel weight quantization method designed to train quantized neural networks (QNNs) to avoid overflow when using low-precision ac…
Quantized Neural Networks for Low-Precision Accumulation with Guaranteed Overflow Avoidance
Ian Colbert, Alessandro Pappalardo, Jakoba Petri-Koenig
We introduce a quantization-aware training algorithm that guarantees avoiding numerical overflow when reducing the precision of accumulators during inference. We leverage weight no…