collaborators

6 papers

cs.DC2026

Empirical Analysis of GPU Frequency Behavior Under ML Workloads

Truong-Thanh Le, Hoang-Loc La, Amir Taherkordi +3

This work presents ongoing research on the frequency scaling behavior of NVIDIA GPUs when executing ML/AI workloads. Our preliminary findings show that, on lower-performance GPUs,…

cs.AI2026

Joint Structural Pruning and Mixed-Precision Quantization for LLM Compression

Hoang-Loc La, Truong-Thanh Le, Amir Taherkordi +1

Recently, the efficiency of Large Language Models (LLMs) deployment has become a critical concern in practical applications. While post-training quantization (PTQ) and structural p…

cs.LG2026

LLM Compression with Jointly Optimizing Architectural and Quantization choices

Hoang-Loc La, Truong-Thanh Le, Amir Taherkordi +1

Deploying large language models (LLMs) is challenging due to their significant memory and computational requirements. While some methods address this by developing small or tiny la…

cs.DC2026

E2LLM: Towards Efficient LLM Serving in Heterogeneous Edge/Fog Environments

Truong-Thanh Le, Amir Taherkordi, Hoang-Loc La +3

Large Language Models (LLMs) have become integral to modern applications, yet their deployment remains challenging. Beyond executing the models themselves, practical deployment mus…

cs.PF2026

PM2Lat: Highly Accurate and Generalized Prediction of DNN Execution Latency on GPUs

Truong-Thanh Le, Hoang-Loc La, Amir Taherkordi +3

We present PM2Lat, a fast and generalized framework for accurately predicting the latency of deep neural network models on GPUs, with special focus on NVIDIA. Unlike prior methods…

cs.LG2025

Kernel-Level Energy-Efficient Neural Architecture Search for Tabular Dataset

Hoang-Loc La, Phuong Hoai Ha

Many studies estimate energy consumption using proxy metrics like memory usage, FLOPs, and inference latency, with the assumption that reducing these metrics will also lower energy…