296 citations · 547 across the 51 of their papers we have counts for
48 papers · 1 filter
FaCTz: Fast Critical-Point and Topology-Aware GPU Compression for Scientific Vector Fields
Mingze Xia, Yuxiao Li, Sheng Di +6
Error-bounded lossy compression is essential for storing and transferring the vector-field data produced by large-scale scientific simulations. Although it enforces a user-specifie…
Evaluating LLM Coding Agents on SZ-Family Lossy Compression Across Architectures
Changqing Li, Shouwei Gao, Kai Zhao +2
Large language model (LLM) coding agents are increasingly applied to code translation and optimization, yet their effectiveness in performance-critical high-performance computing (…
Not All Errors Are Equal: A Systematic Study of Error Propagation in Large Language Model Inference
Yafan Huang, Sheng Di, Guanpeng Li
Large language models (LLMs) are increasingly integrated into high-performance computing (HPC) workflows, accelerating scientific discovery through diverse perspectives such as cod…
ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload
Ziyue Liu, Zhengyang Wang, Ruijie Zhang +7
Pre-training large language models on massive GPU clusters has made hardware faults routine rather than rare, driving the need for resilient training systems. Yet existing framewor…
SplitFT: An Adaptive Federated Split Learning System For LLMs Fine-Tuning
Yimeng Shan, Zhaorui Zhang, Sheng Di +3
Federated Split Learning has been identified as an efficient approach to address the computational resource constraints of clients in classical federated learning, while guaranteei…
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3
Milan Shah, Sheng Di, Michela Becchi
In recent years, novel AI accelerators have emerged as promising alternatives to GPUs for AI model training and inference. One such accelerator, the Cerebras CS-3, has demonstrated…