35 citations · 36 across the 6 of their papers we have counts for
6 papers · 1 filter
Displacement Is Not Direction: Evaluating Fidelity Metrics for Quantized LLM Deployment
Miloš Nikolić, Ali Hadi Zadeh, Enrique Torres Sanchez +1
Fidelity metrics, such as per-token KL divergence (KLD) against a high-precision reference, are often used in practice as low-cost proxies for benchmark quality. We test this pract…
Schrödinger's FP: Dynamic Adaptation of Floating-Point Containers for Deep Learning Training
Miloš Nikolić, Enrique Torres Sanchez, Jiahui Wang +5
The transfer of tensors from/to memory during neural network training dominates time and energy. To improve energy efficiency and performance, research has been exploring ways to u…
Mokey: Enabling Narrow Fixed-Point Inference for Out-of-the-Box Floating-Point Transformer Models
Ali Hadi Zadeh, Mostafa Mahmoud, Ameer Abdelhadi +1
Increasingly larger and better Transformer models keep advancing state-of-the-art accuracy and capability for Natural Language Processing applications. These models demand more com…
GOBO: Quantizing Attention-Based NLP Models for Low Latency and Energy Efficient Inference
Ali Hadi Zadeh, Isak Edo, Omar Mohamed Awad +1
Attention-based models have demonstrated remarkable success in various natural language understanding tasks. However, efficient execution remains a challenge for these models which…
BitPruning: Learning Bitlengths for Aggressive and Accurate Quantization
Miloš Nikolić, Ghouthi Boukli Hacene, Ciaran Bannon +5
Neural networks have demonstrably achieved state-of-the art accuracy using low-bitlength integer quantization, yielding both execution time and energy benefits on existing hardware…
Bit-pragmatic Deep Neural Network Computing
J. Albericio, P. Judd, A. Delmás +2
We quantify a source of ineffectual computations when processing the multiplications of the convolutional layers in Deep Neural Networks (DNNs) and propose Pragmatic (PRA), an arch…