edge llm inference 1energy-efficient computing 1gpu-npu co-execution 1heterogeneous scheduling 1micro-batching 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.DC2026
HeteroMosaic: Exposing and Exploiting Heterogeneous Execution Opportunities for Energy-Efficient Edge LLM Inference
Gregory Hyegang Jun, Wesley Pang, Eddie Richter +4
The paper introduces HeteroMosaic, a scheduling framework that coordinates CPUs, integrated GPUs, and NPUs on edge SoCs to run large language model inference more quickly and with…
cs.DC2026
TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs
Wesley Pang, Gregory Hyegang Jun, Feiyang Liu +1
With the growing demand for on-device LLM inference, edge SoCs increasingly integrate NPUs to improve performance and energy efficiency under tight power and thermal budgets. Howev…
cs.DC2025
Can Asymmetric Tile Buffering Be Beneficial?
Chengyue Wang, Wesley Pang, Xinrui Wu +9
General matrix multiplication (GEMM) is the computational backbone of modern AI workloads, and its efficiency is critically dependent on effective tiling strategies. Conventional a…