1 paper · 1 filter
Mateusz Nowak, Xavier Cadet, Peter Chin
Multiple-choice question (MCQ) benchmarks have been a standard evaluation practice for measuring LLMs' ability to reason and answer knowledge-based questions. Through a synthetic N…