2 papers
cs.AI2024
Inference Performance Optimization for Large Language Models on CPUs
Pujiang He, Shan Zhou, Wenhuan Huang +7
Large language models (LLMs) have shown exceptional performance and vast potential across diverse tasks. However, the deployment of LLMs with high performance in low-resource envir…
cs.DC2024
Distributed Inference Performance Optimization for LLMs on CPUs
Pujiang He, Shan Zhou, Changqing Li +5
Large language models (LLMs) hold tremendous potential for addressing numerous real-world challenges, yet they typically demand significant computational resources and memory. Depl…