1 paper
Qingyuan Li, Ran Meng, Yiduo Li +6
The large language model era urges faster and less costly inference. Prior model compression works on LLMs tend to undertake a software-centric approach primarily focused on the si…