1 paper
Deepanshu Pandey, Arnav Chavan, Nahush Lele +2
Quantization of Large Language Models (LLMs) is often hindered by the sensitivity of the self-attention mechanism to discretization errors. We identify the softmax operator as a bo…