1 paper
Sergio Servantez, Sarah B. Lawsky, Rajiv Jain +2
Reasoning benchmarks have played a crucial role in the progress of language models. Yet rigorous evaluation remains a significant challenge as static question-answer pairs provide…