1 paper
Binwen Tan, Jingchao Wang, Dengzhe Hou +6
Reinforcement learning post-training unlocks complex reasoning in LLMs. Yet benchmark scores reveal only whether a model improved, not what changed inside it, nor how it splits fin…