2 papers
cs.LG2025
Understanding Silent Data Corruption in LLM Training
Jeffrey Ma, Hengzhi Pei, Leonard Lausen +1
As the scale of training large language models (LLMs) increases, one emergent failure is silent data corruption (SDC), where hardware produces incorrect computations without explic…
cs.CL2024
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Boxin Wang, Weixin Chen, Hengzhi Pei +16
Generative Pre-trained Transformer (GPT) models have exhibited exciting progress in their capabilities, capturing the interest of practitioners and the public alike. Yet, while the…