1 paper
Haoyu Qiao, Hao Zhang, Shanwen Mao +2
Large language models (LLMs) deliver impressive capabilities but incur substantial inference latency and cost, which hinders their deployment in latency-sensitive and resource-cons…