7 citations · 7 across the 6 of their papers we have counts for
1 paper · 1 filter
Sam Bowyer, Laurence Aitchison, Desi R. Ivanova
Rigorous statistical evaluations of large language models (LLMs), including valid error bars and significance testing, are essential for meaningful and reliable performance assessm…