7 citations · 7 across the 1 of their papers we have counts for
1 paper
Yifan Song, Guoyin Wang, Sujian Li +1
Current evaluations of large language models (LLMs) often overlook non-determinism, typically focusing on a single output per example. This limits our understanding of LLM performa…