1 paper
Ying Gu, Mei Chee Leong, Hui Li Tan +3
Dominant accuracy evaluation might reward unwarranted guessing of Large Language Models, and it might not be applicable to novel tasks for model validation without ground-truth (gt…