1 paper · 1 filter
Wenxiang Lin, Xinglin Pan, Shaohuai Shi +2
Large language models~(LLMs) are known for their high demand on computing resources and memory due to their substantial model size, which leads to inefficient inference on moderate…