3 papers
cs.DC2026
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
Ruibo Fan, Xiangrui Yu, Xinglin Pan +5
Lossless model compression holds tremendous promise for alleviating the memory and bandwidth bottlenecks in bit-exact Large Language Model (LLM) serving. However, existing approach…
cs.DC2025
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
Weile Luo, Ruibo Fan, Zeyu Li +4
This study presents a comprehensive multi-level analysis of the NVIDIA Hopper GPU architecture, focusing on its performance characteristics and novel features. We benchmark Hopper'…
cs.PF2024
DSO: A GPU Energy Efficiency Optimizer by Fusing Dynamic and Static Information
Qiang Wang, Laiyi Li, Weile Luo +2
Increased reliance on graphics processing units (GPUs) for high-intensity computing tasks raises challenges regarding energy consumption. To address this issue, dynamic voltage and…