1 paper
Crish Nagarkar, Leonid Bogachev, Serge Sharoff
This paper investigates the ability of large language models (LLMs) to solve statistical tasks, as well as their capacity to assess the quality of reasoning. While state-of-the-art…