4 papers
Structured Scaling of AI Discovery Across Diverse Scientific Domains
Haotian Ye, Haowei Lin, Jingyi Tang +30
Scientific discovery often requires many cycles of proposing, testing, and refining candidate solutions. Language models can increasingly participate in these loops, but simply gen…
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
Ruibo Fan, Xiangrui Yu, Xinglin Pan +5
Lossless model compression holds tremendous promise for alleviating the memory and bandwidth bottlenecks in bit-exact Large Language Model (LLM) serving. However, existing approach…
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
Weile Luo, Ruibo Fan, Zeyu Li +4
This study presents a comprehensive multi-level analysis of the NVIDIA Hopper GPU architecture, focusing on its performance characteristics and novel features. We benchmark Hopper'…
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
Yuxin Wang, Yuhan Chen, Zeyu Li +11
Serving systems for Large Language Models (LLMs) are often optimized to improve quality of service (QoS) and throughput. However, due to the lack of open-source LLM serving workloa…