1 paper
Xinlin Li, Timothy Chou, Josh Fromm +3
Post-training weight quantization is crucial for reducing the memory and inference cost of large language models (LLMs), yet pushing the average precision below 4 bits remains chal…