3 papers
cs.AR2026
The Anatomy of Silent Data Corruption: GPU Error Pattern Study and Modeling Guidance
Chung-Hsuan Tung, Yanxiang Huang, Nirmal Saxena +5
Silent data corruption (SDC) threatens the reliability of large-scale GPU clusters used for training large language models, yet its rarity and lack of explicit error signals make a…
cs.AR2026
LLM-PRISM: Characterizing Silent Data Corruption from Permanent GPU Faults in LLM Training
Abhishek Tyagi, Saurabh Hukerikar, Nirmal Saxena +4
Large-scale LLM training is increasingly susceptible to hardware defects stemming from manufacturing escapes and silicon aging. These defects manifest as Silent Data Corruption (SD…
cs.DC2024
Optimal Checkpoint Interval with Availability as an Objective Function
Nirmal Raj Saxena, Saurabh Hukerikar, Mikolaj Blaz +1
We present a simplified derivation of the optimal checkpoint interval in Young_1974 [1]. The optimal checkpoint interval derivation in [1] is based on minimizing the total lost tim…