148 citations · 171 across the 10 of their papers we have counts for
12 papers
GreenPipe: Power Modeling for Containerized DNN Inference on Kubernetes Edge Nodes
Mengxue Wang, Peini Liu, Amir Taherkordi +1
Distributed DNN inference is increasingly deployed in containerized edge-cloud environments, where workloads run on-device or are exposed to remote clients over the network. Accura…
GAPL: Grounded Action-effect Policy Learning for LLM-Based Trajectory Planning
Zhihong Cui, Hengyu Liu, Zhangkai Wu +5
Trajectory planning for autonomous driving requires both high-level reasoning and precise low-level control. Large Language Models (LLMs) offer semantic-rich planning capabilities,…
Empirical Analysis of GPU Frequency Behavior Under ML Workloads
Truong-Thanh Le, Hoang-Loc La, Amir Taherkordi +3
This work presents ongoing research on the frequency scaling behavior of NVIDIA GPUs when executing ML/AI workloads. Our preliminary findings show that, on lower-performance GPUs,…
Joint Structural Pruning and Mixed-Precision Quantization for LLM Compression
Hoang-Loc La, Truong-Thanh Le, Amir Taherkordi +1
Recently, the efficiency of Large Language Models (LLMs) deployment has become a critical concern in practical applications. While post-training quantization (PTQ) and structural p…
LLM Compression with Jointly Optimizing Architectural and Quantization choices
Hoang-Loc La, Truong-Thanh Le, Amir Taherkordi +1
Deploying large language models (LLMs) is challenging due to their significant memory and computational requirements. While some methods address this by developing small or tiny la…
E2LLM: Towards Efficient LLM Serving in Heterogeneous Edge/Fog Environments
Truong-Thanh Le, Amir Taherkordi, Hoang-Loc La +3
Large Language Models (LLMs) have become integral to modern applications, yet their deployment remains challenging. Beyond executing the models themselves, practical deployment mus…