papers

Publications (7)

cs.AR2024

Benchmarking and Dissecting the Nvidia Hopper GPU Architecture

Weile Luo, Ruibo Fan, Zeyu Li +3

Graphics processing units (GPUs) are continually evolving to cater to the computational demands of contemporary general-purpose workloads, particularly those driven by artificial i…

cs.DC2026

ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression

Ruibo Fan, Xiangrui Yu, Xinglin Pan +5

Lossless model compression holds tremendous promise for alleviating the memory and bandwidth bottlenecks in bit-exact Large Language Model (LLM) serving. However, existing approach…

cs.DC2025

Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis

Weile Luo, Ruibo Fan, Zeyu Li +4

This study presents a comprehensive multi-level analysis of the NVIDIA Hopper GPU architecture, focusing on its performance characteristics and novel features. We benchmark Hopper'…

cs.LG2026

Dissecting Outlier Dynamics in LLM NVFP4 Pretraining

Peijie Dong, Ruibo Fan, Yuechen Tao +11

Training large language models using 4-bit arithmetic enhances throughput and memory efficiency. Yet, the limited dynamic range of FP4 increases sensitivity to outliers. While NVFP…

cs.PF2023

Dissecting the Runtime Performance of the Training, Fine-tuning, and Inference of Large Language Models

Longteng Zhang, Xiang Liu, Zeyu Li +8

Large Language Models (LLMs) have seen great advance in both academia and industry, and their popularity results in numerous open-source frameworks and techniques in accelerating L…

cs.LG2024

STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs

Peijie Dong, Lujun Li, Yuedong Zhong +8

In this paper, we present the first structural binarization method for LLM compression to less than 1-bit precision. Although LLMs have achieved remarkable performance, their memor…