3 papers
cs.DC2026
Empirical Analysis of GPU Frequency Behavior Under ML Workloads
Truong-Thanh Le, Hoang-Loc La, Amir Taherkordi +3
This work presents ongoing research on the frequency scaling behavior of NVIDIA GPUs when executing ML/AI workloads. Our preliminary findings show that, on lower-performance GPUs,…
cs.DC2026
E2LLM: Towards Efficient LLM Serving in Heterogeneous Edge/Fog Environments
Truong-Thanh Le, Amir Taherkordi, Hoang-Loc La +3
Large Language Models (LLMs) have become integral to modern applications, yet their deployment remains challenging. Beyond executing the models themselves, practical deployment mus…
cs.PF2026
PM2Lat: Highly Accurate and Generalized Prediction of DNN Execution Latency on GPUs
Truong-Thanh Le, Hoang-Loc La, Amir Taherkordi +3
We present PM2Lat, a fast and generalized framework for accurately predicting the latency of deep neural network models on GPUs, with special focus on NVIDIA. Unlike prior methods…