activity
20132024
most citedA White Paper on Neural Network Quantization

13 citations · 18 across the 4 of their papers we have counts for

collaborators

10 papers

cs.CL20241 cited

Context-Aware Clustering using Large Language Models

Sindhu Tipirneni, Ravinarayana Adkathimar, Nurendra Choudhary +5

Despite the remarkable success of Large Language Models (LLMs) in text understanding and generation, their potential for text clustering tasks remains underexplored. We observed th…

cs.LG202113 cited

A White Paper on Neural Network Quantization

Markus Nagel, Marios Fournarakis, Rana Ali Amjad +3

While neural networks have advanced the frontiers in many applications, they often come at a high computational cost. Reducing the power and latency of neural network inference is…

cs.IT2020

Invertible Low-Divergence Coding

Patrick Schulte, Rana Ali Amjad, Thomas Wiegart +1

Several applications in communication, control, and learning require approximating target distributions to within small informational divergence (I-divergence). The additional requ…

cs.LG2020

Bayesian Bits: Unifying Quantization and Pruning

Mart van Baalen, Christos Louizos, Markus Nagel +4

We introduce Bayesian Bits, a practical method for joint mixed precision quantization and pruning through gradient based optimization. Bayesian Bits employs a novel decomposition o…

cs.LG2020

Up or Down? Adaptive Rounding for Post-Training Quantization

Markus Nagel, Rana Ali Amjad, Mart van Baalen +2

When quantizing neural networks, assigning each floating-point weight to its nearest fixed-point value is the predominant approach. We find that, perhaps surprisingly, this is not…

cs.LG2019

Class-Conditional Compression and Disentanglement: Bridging the Gap between Neural Networks and Naive Bayes Classifiers

Rana Ali Amjad, Bernhard C. Geiger

In this draft, which reports on work in progress, we 1) adapt the information bottleneck functional by replacing the compression term by class-conditional compression, 2) relax thi…