5 citations · 5 across the 4 of their papers we have counts for
1 paper · 1 filter
Jiamin Chen, Yidi Wu, Qiexiang Wang +6
Widely used language-model benchmarks are increasingly saturated, with frontier systems often receiving near-tied scores that standard metrics cannot resolve. Rather than construct…