1 paper
Xinyuan Li, Murong Xu, Wenbiao Tao +4
Large language models (LLMs) achieve high performance on mathematical reasoning, but these results can be inflated by training data leakage or superficial pattern matching rather t…