Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference
Jiayi Yuan, Hao Li, Xinheng Ding +7
Large Language Models (LLMs) are now integral across various domains and have demonstrated impressive performance. Progress, however, rests on the premise that benchmark scores are…
cs.CL2025
The Science of Evaluating Foundation Models
Jiayi Yuan, Jiamu Zhang, Andrew Wen +1
The emergent phenomena of large foundation models have revolutionized natural language processing. However, evaluating these models presents significant challenges due to their siz…