29 papers · 1 filter
FaCTz: Fast Critical-Point and Topology-Aware GPU Compression for Scientific Vector Fields
Mingze Xia, Yuxiao Li, Sheng Di +6
Error-bounded lossy compression is essential for storing and transferring the vector-field data produced by large-scale scientific simulations. Although it enforces a user-specifie…
Not All Errors Are Equal: A Systematic Study of Error Propagation in Large Language Model Inference
Yafan Huang, Sheng Di, Guanpeng Li
Large language models (LLMs) are increasingly integrated into high-performance computing (HPC) workflows, accelerating scientific discovery through diverse perspectives such as cod…
SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs
Jin Lee, Zhonghao Chen, Xuhang He +6
In large-scale LLM pre-training systems with 100k+ GPUs, failures become the norm rather than the exception, and restart costs can dominate wall-clock training time. However, exist…
ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload
Ziyue Liu, Zhengyang Wang, Ruijie Zhang +7
Pre-training large language models on massive GPU clusters has made hardware faults routine rather than rare, driving the need for resilient training systems. Yet existing framewor…
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3
Milan Shah, Sheng Di, Michela Becchi
In recent years, novel AI accelerators have emerged as promising alternatives to GPUs for AI model training and inference. One such accelerator, the Cerebras CS-3, has demonstrated…
SplitFT: An Adaptive Federated Split Learning System For LLMs Fine-Tuning
Yimeng Shan, Zhaorui Zhang, Sheng Di +3
Federated Split Learning has been identified as an efficient approach to address the computational resource constraints of clients in classical federated learning, while guaranteei…