1 paper
Hwiyeol Jo, Joosung Lee, Jaehone Lee +3
Evaluating generative models, such as large language models (LLMs), commonly involves question-answering tasks where the final answer is selected based on probability of answer cho…