238 citations · 448 across the 4 of their papers we have counts for
1 paper · 1 filter
Pranav Nair, Puranjay Datta, Jeff Dean +2
Quantizing model weights is critical for reducing the communication and inference costs of large models. However, quantizing models -- especially to low precisions like int4 or int…