1 paper
Arda Yüksel, Abdullatif Köksal, Lütfi Kerem Şenel +2
Multiple choice question answering tasks evaluate the reasoning, comprehension, and mathematical abilities of Large Language Models (LLMs). While existing benchmarks employ automat…