1 paper
Simon Frieder, Jonas Bayer, Sam Looi +13
The datasets and benchmarks commonly used to train and evaluate the mathematical capabilities of AI-based mathematical copilots (primarily large language models) exhibit several sh…