1 paper · 1 filter
Nikhil Chandak, Shashwat Goel, Ameya Prabhu +2
Multiple choice benchmarks have long been the workhorse of language model evaluation because grading multiple choice is objective and easy to automate. However, we show multiple ch…