2 papers
cs.DC2026
A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM
Shaoke Xi, ChonLam Lao, Boyi Jia +11
Large language model (LLM) training today runs on clusters spanning thousands of GPUs. While this scale enables rapid model advances, developing, debugging, and performance-tuning…
cs.NI2024
Accelerating Stateful Network Applications with Performance Prediction on SoC SmartNICs
Shaoke Xi, Jiaqi Gao, Mengqi Liu +7
Offloading stateful network functions to multi-threaded SoC SmartNICs promises significant performance and cost benefits. However, realizing this potential is hindered by two funda…