1 paper
Minyang Tian, Luyu Gao, Shizhuo Dylan Zhang +27
Since language models (LMs) now outperform average humans on many challenging tasks, it has become increasingly difficult to develop challenging, high-quality, and realistic evalua…