1 paper
Yiyang Li, Yonghuang Wu, Ying Luo +5
Evaluations of large language models (LLMs) suffer from instability, where small changes of random factors such as few-shot examples can lead to drastic fluctuations of scores and…