4 papers
SocraticPO: Policy Optimization via Interactive Guidance
Zirui Liu, Jie Ouyang, Qi Liu +8
Reinforcement learning (RL) for large language models usually supervises reasoning with scalar outcome rewards, such as binary correctness. Such rewards provide an optimization dir…
CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive Perspective
Jiayu Liu, Zhenya Huang, Wei Dai +7
Although large language models (LLMs) show promise in solving complex mathematical tasks, existing evaluation paradigms rely solely on a coarse measure of overall answer accuracy,…
MMATH: A Multilingual Benchmark for Mathematical Reasoning
Wenyang Luo, Wayne Xin Zhao, Jing Sha +2
The advent of large reasoning models, such as OpenAI o1 and DeepSeek R1, has significantly advanced complex reasoning tasks. However, their capabilities in multilingual complex rea…
DACO: Towards Application-Driven and Comprehensive Data Analysis via Code Generation
Xueqing Wu, Rui Zheng, Jingzhen Sha +6
Data analysis is a crucial analytical process to generate in-depth studies and conclusive insights to comprehensively answer a given user query for tabular data. In this work, we a…