4 citations · 4 across the 12 of their papers we have counts for
1 paper · 2 filters
Mateusz Nowak, Xavier Cadet, Peter Chin
Multiple-choice question (MCQ) benchmarks have been a standard evaluation practice for measuring LLMs' ability to reason and answer knowledge-based questions. Through a synthetic N…