1 paper
Hongyao Liu, Liuqun Zhai, Junyi Wang +1
Efficient inference for on-device Large Language Models (LLMs) remains challenging due to limited hardware resources and the high cost of the prefill stage, which processes the ful…