3 papers
cs.DC2026
Empirical Analysis of GPU Frequency Behavior Under ML Workloads
Truong-Thanh Le, Hoang-Loc La, Amir Taherkordi +3
This work presents ongoing research on the frequency scaling behavior of NVIDIA GPUs when executing ML/AI workloads. Our preliminary findings show that, on lower-performance GPUs,…
cs.AI2026
Joint Structural Pruning and Mixed-Precision Quantization for LLM Compression
Hoang-Loc La, Truong-Thanh Le, Amir Taherkordi +1
Recently, the efficiency of Large Language Models (LLMs) deployment has become a critical concern in practical applications. While post-training quantization (PTQ) and structural p…
cs.LG2026
LLM Compression with Jointly Optimizing Architectural and Quantization choices
Hoang-Loc La, Truong-Thanh Le, Amir Taherkordi +1
Deploying large language models (LLMs) is challenging due to their significant memory and computational requirements. While some methods address this by developing small or tiny la…