1 paper · 1 filter
Yunlong Deng, Boyang Sun, Yan Li +4
Due to their inherent complexity, reasoning tasks have long been regarded as rigorous benchmarks for assessing the capabilities of machine learning models, especially large languag…