1 paper
Haocheng Lu, Minjun Zhu, Henry Yu
Large language models (LLMs) continue to struggle with mathematical reasoning, and common post-training pipelines often reduce each generated solution to a binary outcome: correct…