1 paper · 1 filter
Jhen-Ke Lin
Language models are commonly compared by averaging scores across a benchmark list with equal weight. Such lists grow through publication outside an explicit measurement design, so…