1 paper · 1 filter
Shreyas K Chandrahas
The dominant practice in language model evaluation is to report a single accuracy number per model and declare the higher one better, without testing whether the gap could plausibl…