1 paper · 1 filter
Kaichen Zhang, Bo Li, Peiyuan Zhang +8
The advances of large foundation models necessitate wide-coverage, low-cost, and zero-contamination benchmarks. Despite continuous exploration of language model evaluations, compre…