2 papers
cs.LG2026
TrainSDC: Characterizing and Mitigating Silent Data Corruption in Large Language Model Training
Zhipeng Xia, Haotian Xu, Siyu Yun +4
LLM training is increasingly vulnerable to silent data corruption (SDC), yet existing protection methods largely treat Transformer computations uniformly because their vulnerabilit…
cs.LG2026
ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning
Wenwu Fan, Qihong Lin, Zhijie Xia +4
Reinforcement Learning (RL) training for Large Language Models (LLMs) often suffers from instability due to the discrepancy between training and inference. This training-inference…