1 paper
Bianca Raimondi, Francesco Pivi, Davide Evangelista +1
The evaluation of Large Language Models (LLMs) on mathematical reasoning has largely focused on elementary problems, competition-style questions, or formal theorem proving, leaving…