activity
20162026
most citedMokey: Enabling Narrow Fixed-Point Inference for Out-of-the-Box Floating-Point Transformer Models

35 citations · 36 across the 6 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2026

Displacement Is Not Direction: Evaluating Fidelity Metrics for Quantized LLM Deployment

Miloš Nikolić, Ali Hadi Zadeh, Enrique Torres Sanchez +1

Fidelity metrics, such as per-token KL divergence (KLD) against a high-precision reference, are often used in practice as low-cost proxies for benchmark quality. We test this pract…

cs.LG2022

Schrödinger's FP: Dynamic Adaptation of Floating-Point Containers for Deep Learning Training

Miloš Nikolić, Enrique Torres Sanchez, Jiahui Wang +5

The transfer of tensors from/to memory during neural network training dominates time and energy. To improve energy efficiency and performance, research has been exploring ways to u…

cs.LG2022★ 35 cited

Mokey: Enabling Narrow Fixed-Point Inference for Out-of-the-Box Floating-Point Transformer Models

Ali Hadi Zadeh, Mostafa Mahmoud, Ameer Abdelhadi +1

Increasingly larger and better Transformer models keep advancing state-of-the-art accuracy and capability for Natural Language Processing applications. These models demand more com…

cs.LG2020

GOBO: Quantizing Attention-Based NLP Models for Low Latency and Energy Efficient Inference

Ali Hadi Zadeh, Isak Edo, Omar Mohamed Awad +1

Attention-based models have demonstrated remarkable success in various natural language understanding tasks. However, efficient execution remains a challenge for these models which…

cs.LG2020

BitPruning: Learning Bitlengths for Aggressive and Accurate Quantization

Miloš Nikolić, Ghouthi Boukli Hacene, Ciaran Bannon +5

Neural networks have demonstrably achieved state-of-the art accuracy using low-bitlength integer quantization, yielding both execution time and energy benefits on existing hardware…

cs.LG2016

Bit-pragmatic Deep Neural Network Computing

J. Albericio, P. Judd, A. Delmás +2

We quantify a source of ineffectual computations when processing the multiplications of the convolutional layers in Deep Neural Networks (DNNs) and propose Pragmatic (PRA), an arch…