2 papers
cs.LG2026
SocraticPO: Policy Optimization via Interactive Guidance
Zirui Liu, Jie Ouyang, Qi Liu +8
Reinforcement learning (RL) for large language models usually supervises reasoning with scalar outcome rewards, such as binary correctness. Such rewards provide an optimization dir…
cs.CL2025
MMATH: A Multilingual Benchmark for Mathematical Reasoning
Wenyang Luo, Wayne Xin Zhao, Jing Sha +2
The advent of large reasoning models, such as OpenAI o1 and DeepSeek R1, has significantly advanced complex reasoning tasks. However, their capabilities in multilingual complex rea…