1 paper
Suchit Gupte, Vishnu Kabir Chhabra, Mohammad Mahdi Khalili
Modern LLMs face inference efficiency challenges due to their scale. To address this, many compression methods have been proposed, such as pruning and quantization. However, the ef…