5 papers
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
Wangsong Yin, Daliang Xu, Mengwei Xu +2
On-device running Large Language Models (LLMs) is nowadays a critical enabler towards preserving user privacy. We observe that the attention operator falls back from the special-pu…
Elastic On-Device LLM Service
Wangsong Yin, Rongjie Yi, Daliang Xu +3
On-device Large Language Models (LLMs) are transforming mobile AI, catalyzing applications like UI automation without privacy concerns. Nowadays the common practice is to deploy a…
SCOPE: Performance Testing for Serverless Computing
Jinfeng Wen, Zhenpeng Chen, Jianshu Zhao +5
Serverless computing is a popular cloud computing paradigm that has found widespread adoption across various online workloads. It allows software engineers to develop cloud applica…
Unveiling Overlooked Performance Variance in Serverless Computing
Jinfeng Wen, Zhenpeng Chen, Federica Sarro +1
Serverless computing is an emerging cloud computing paradigm for developing applications at the function level, known as serverless functions. Due to the highly dynamic execution e…
Fast On-device LLM Inference with NPUs
Daliang Xu, Hao Zhang, Liming Yang +4
On-device inference for Large Language Models (LLMs), driven by increasing privacy concerns and advancements of mobile-sized models, has gained significant interest. However, even…