papers
Publications (8)
cs.CV2018
FINN-L: Library Extensions and Design Trade-off Analysis for Variable Precision LSTM Networks on FPGAs
Vladimir Rybalkin, Alessandro Pappalardo, Muhammad Mohsin Ghaffar +3
cs.LG2021
Ps and Qs: Quantization-aware pruning for efficient low latency neural network inference
Benjamin Hawks, Javier Duarte, Nicholas J. Fraser +3
cs.CV2024
Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs
Shivam Aggarwal, Hans Jakob Damsgaard, Alessandro Pappalardo +4
cs.LG2024
A2Q+: Improving Accumulator-Aware Weight Quantization
Ian Colbert, Alessandro Pappalardo, Jakoba Petri-Koenig +1
cs.LG2022
QONNX: Representing Arbitrary-Precision Quantized Neural Networks
Alessandro Pappalardo, Yaman Umuroglu, Michaela Blott +11
cs.LG2023
A2Q: Accumulator-Aware Quantization with Guaranteed Overflow Avoidance
Ian Colbert, Alessandro Pappalardo, Jakoba Petri-Koenig
cs.LG2025
Apple Intelligence Foundation Language Models: Tech Report 2025
Ethan Li, Anders Boesen Lindbo Larsen, Chen Zhang +395
cs.LG2023
Quantized Neural Networks for Low-Precision Accumulation with Guaranteed Overflow Avoidance
Ian Colbert, Alessandro Pappalardo, Jakoba Petri-Koenig