Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Keyword Matters: Unveiling the Energy Sensitivity of On-Device LLM Prompting
Ruiyi Tao, Xiaolong Tu, Haoxin Wang
Large Language Models (LLMs) are increasingly deployed on mobile and embedded devices to improve privacy and reduce network latency. Yet on-device inference faces a fundamental con…
cs.AI2026
Transferable Latency Prediction for Fast LLM Screening on Heterogeneous Edge Devices
Xiaolong Tu, Vinod K. Mishra, Venkat R. Dasari +2
Accurate latency prediction is critical for deploying large language models (LLMs) on heterogeneous edge devices, where inference latency is affected by model architecture, prompt…