edge llm inference 1energy-efficient computing 1gpu-npu co-execution 1heterogeneous scheduling 1micro-batching 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.DC2026
HeteroMosaic: Exposing and Exploiting Heterogeneous Execution Opportunities for Energy-Efficient Edge LLM Inference
Gregory Hyegang Jun, Wesley Pang, Eddie Richter +4
The paper introduces HeteroMosaic, a scheduling framework that coordinates CPUs, integrated GPUs, and NPUs on edge SoCs to run large language model inference more quickly and with…
cs.DC2026
TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs
Wesley Pang, Gregory Hyegang Jun, Feiyang Liu +1
With the growing demand for on-device LLM inference, edge SoCs increasingly integrate NPUs to improve performance and energy efficiency under tight power and thermal budgets. Howev…