8 papers
Beyond the Sampled Token: Preserving Candidate Support in RLVR
Ruotian Peng, Yi Ren, Zhouliang Yu +2
We revisit exploration collapse in reinforcement learning with verifiable rewards (RLVR), from the perspective of the \emph{candidate distribution} for next-token prediction. We fo…
Sample-Efficient Post-Training for LEGO Spatial-Physics Reasoning
Yuhuan Yuan, Zhouliang Yu, Minghao Liu +2
LLM-based LEGO assembly requires both semantic grounding and physical feasibility. In this paper, we identify a data-induced failure mode, physhack, in which generated assemblies s…
PAPO: Stabilizing Rubric Integration Training via Decoupled Advantage Normalization
Zelin Tan, Zhouliang Yu, Bohan Lin +9
We propose Process-Aware Policy Optimization (PAPO), a method that integrates process-level evaluation into Group Relative Policy Optimization (GRPO) through decoupled advantage no…
CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization
Zhongyuan Peng, Yifan Yao, Kaijing Ma +16
Translating natural language mathematical statements into formal, executable code is a fundamental challenge in automated theorem proving. While prior work has focused on generatio…
Generating Symbolic World Models via Test-time Scaling of Large Language Models
Zhouliang Yu, Yuhuan Yuan, Tim Z. Xiao +5
Solving complex planning problems requires Large Language Models (LLMs) to explicitly model the state transition to avoid rule violations, comply with constraints, and ensure optim…
FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models
Zhouliang Yu, Ruotian Peng, Keyi Ding +10
Formal mathematical reasoning remains a critical challenge for artificial intelligence, hindered by limitations of existing benchmarks in scope and scale. To address this, we prese…