3 citations · 3 across the 3 of their papers we have counts for
1 paper · 1 filter
Qingyuan Li, Ran Meng, Yiduo Li +6
The large language model era urges faster and less costly inference. Prior model compression works on LLMs tend to undertake a software-centric approach primarily focused on the si…