1 paper
Utkarsh Saxena, Sayeh Sharify, Kaushik Roy +1
Post-training quantization (PTQ) of large language models (LLMs) holds the promise in reducing the prohibitive computational cost at inference time. Quantization of all weight, act…