1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Chengyu Shen, Yanheng Hou, Minghui Pan +8
Reliable evaluation is essential for developing and deploying large language models, yet in practice it often requires substantial manual effort: practitioners must identify approp…