11 papers · 1 filter
Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals
Sirui Chen, Lei Xu, Yuying Zhao +6
Recent RL methods have substantially improved the reasoning abilities of LLMs. Existing reward designs mainly follow two paradigms: (1) Reinforcement learning with verifiable rewar…
Code as Agent Harness
Xuying Ning, Katherine Tieu, Dongqi Fu +39
Recent large language models (LLMs) have demonstrated strong capabilities in understanding and generating code, from competitive programming to repository-level software engineerin…
Can Post-Training Transform LLMs into Causal Reasoners?
Junqi Chen, Sirui Chen, Chaochao Lu
Causal inference is essential for decision-making but remains challenging for non-experts. While large language models (LLMs) show promise in this domain, their precise causal esti…
CauScientist: Teaching LLMs to Respect Data for Causal Discovery
Bo Peng, Sirui Chen, Lei Xu +1
Causal discovery is fundamental to scientific understanding and reliable decision-making. Existing approaches face critical limitations: purely data-driven methods suffer from stat…
DEPO: Dual-Efficiency Preference Optimization for LLM Agents
Sirui Chen, Mengshi Zhao, Lei Xu +5
Recent advances in large language models (LLMs) have greatly improved their reasoning and decision-making abilities when deployed as agents. Richer reasoning, however, often comes…
Synthesis by Design: Controlled Data Generation via Structural Guidance
Lei Xu, Sirui Chen, Yuxuan Huang +1
Mathematical reasoning remains challenging for LLMs due to complex logic and the need for precise computation. Existing methods enhance LLM reasoning by synthesizing datasets throu…