Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Diffusion-Proof: Recipe for Formal Theorem Proving Beyond Auto-Regressive Generation
Ruida Wang, Rui Pan, Pengcheng Wang +2
Enhancing the formal math reasoning capabilities of Large Language Models (LLMs) has become a key focus in both mathematical and computer science communities in recent years. While…
cs.LG2026
Future-KL Regularized GRPO: Process-Level Credit Assignment from -Divergence Regularization
Jiarui Yao, Ruida Wang, Hao Bai +1
Group Relative Policy Optimization (GRPO) is widely used for critic-free Large Language Model (LLM) post-training, but its KL regularization is usually implemented as a local loss-…
cs.LG2026
GAR: Generative Adversarial Reinforcement Learning for Formal Theorem Proving
Ruida Wang, Jiarui Yao, Rui Pan +2
Solving math problems through verifiable languages such as Lean has significantly impacted both the mathematics and computer science communities. Current state-of-the-art models ar…