7 citations · 9 across the 7 of their papers we have counts for
1 paper · 2 filters
Aryo Pradipta Gema, Joshua Ong Jun Leang, Giwon Hong +13
Maybe not. We identify and analyse errors in the popular Massive Multitask Language Understanding (MMLU) benchmark. Even though MMLU is widely adopted, our analysis demonstrates nu…