1 paper
LucÃa M. Cabrera, Isaac Saxton-Knight, Jocelyn D'Arcy
Mathematical benchmarks consisting of a range of mathematics problems are widely used to evaluate the reasoning abilities of large language models, yet little is known about how th…