4 papers
The Anatomy of Silent Data Corruption: GPU Error Pattern Study and Modeling Guidance
Chung-Hsuan Tung, Yanxiang Huang, Nirmal Saxena +5
Silent data corruption (SDC) threatens the reliability of large-scale GPU clusters used for training large language models, yet its rarity and lack of explicit error signals make a…
LLM-PRISM: Characterizing Silent Data Corruption from Permanent GPU Faults in LLM Training
Abhishek Tyagi, Saurabh Hukerikar, Nirmal Saxena +4
Large-scale LLM training is increasingly susceptible to hardware defects stemming from manufacturing escapes and silicon aging. These defects manifest as Silent Data Corruption (SD…
SHUFFLESPARSE: Learned Shuffles for Structured Sparse Networks
Abhishek Tyagi, Arjun Iyer, Liam Young +3
Structured weight sparsity accelerates training and inference on modern GPUs, but it trails unstructured dynamic sparse training (DST) in accuracy especially at extreme sparsity. W…
Dynamic Sparse Training of Diagonally Sparse Networks
Abhishek Tyagi, Arjun Iyer, William H Renninger +2
Recent advances in Dynamic Sparse Training (DST) have pushed the frontier of sparse neural network training in structured and unstructured contexts, matching dense-model performanc…