4 papers
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
Wangsong Yin, Daliang Xu, Mengwei Xu +2
On-device running Large Language Models (LLMs) is nowadays a critical enabler towards preserving user privacy. We observe that the attention operator falls back from the special-pu…
Elastic On-Device LLM Service
Wangsong Yin, Rongjie Yi, Daliang Xu +3
On-device Large Language Models (LLMs) are transforming mobile AI, catalyzing applications like UI automation without privacy concerns. Nowadays the common practice is to deploy a…
Fast On-device LLM Inference with NPUs
Daliang Xu, Hao Zhang, Liming Yang +4
On-device inference for Large Language Models (LLMs), driven by increasing privacy concerns and advancements of mobile-sized models, has gained significant interest. However, even…
Research on WebAssembly Runtimes: A Survey
Yixuan Zhang, Mugeng Liu, Haoyu Wang +3
WebAssembly (abbreviated as Wasm) was initially introduced for the Web but quickly extended its reach into various domains beyond the Web. To create Wasm applications, developers c…