2 papers
cs.DC2026
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
Ruibo Fan, Xiangrui Yu, Xinglin Pan +5
Lossless model compression holds tremendous promise for alleviating the memory and bandwidth bottlenecks in bit-exact Large Language Model (LLM) serving. However, existing approach…
cs.DC2025
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
Weile Luo, Ruibo Fan, Zeyu Li +4
This study presents a comprehensive multi-level analysis of the NVIDIA Hopper GPU architecture, focusing on its performance characteristics and novel features. We benchmark Hopper'…