5 papers
Keyword Matters: Unveiling the Energy Sensitivity of On-Device LLM Prompting
Ruiyi Tao, Xiaolong Tu, Haoxin Wang
Large Language Models (LLMs) are increasingly deployed on mobile and embedded devices to improve privacy and reduce network latency. Yet on-device inference faces a fundamental con…
Transferable Latency Prediction for Fast LLM Screening on Heterogeneous Edge Devices
Xiaolong Tu, Vinod K. Mishra, Venkat R. Dasari +2
Accurate latency prediction is critical for deploying large language models (LLMs) on heterogeneous edge devices, where inference latency is affected by model architecture, prompt…
PlatformX: An End-to-End Transferable Platform for Energy-Efficient Neural Architecture Search
Xiaolong Tu, Dawei Chen, Kyungtae Han +2
Hardware-Aware Neural Architecture Search (HW-NAS) has emerged as a powerful tool for designing efficient deep neural networks (DNNs) tailored to edge devices. However, existing me…
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
Haoxin Wang, Xiaolong Tu, Hongyu Ke +3
Large Language Models (LLMs) are increasingly integrated into everyday applications, but their prevalent cloud-based deployment raises growing concerns around data privacy and long…
GreenAuto: An Automated Platform for Sustainable AI Model Design on Edge Devices
Xiaolong Tu, Dawei Chen, Kyungtae Han +2
We present GreenAuto, an end-to-end automated platform designed for sustainable AI model exploration, generation, deployment, and evaluation. GreenAuto employs a Pareto front-based…