1 paper · 1 filter
Igor Ivanov, Dmitrii Volkov
Recent work showed that small changes in benchmark questions can reduce LLMs' reasoning and recall. We explore two such changes: pairing questions and adding more answer options, o…