1 paper · 1 filter
Xinhao Qu, Qiang Heng, Hao Zeng +1
Evaluation of large language models (LLMs) is increasingly critical, yet standard benchmarking methods rely on average accuracy, overlooking both the inherent stochasticity of LLM…