Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Beyond the Sampled Token: Preserving Candidate Support in RLVR
Ruotian Peng, Yi Ren, Zhouliang Yu +2
We revisit exploration collapse in reinforcement learning with verifiable rewards (RLVR), from the perspective of the \emph{candidate distribution} for next-token prediction. We fo…
cs.AI2025
FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models
Zhouliang Yu, Ruotian Peng, Keyi Ding +10
Formal mathematical reasoning remains a critical challenge for artificial intelligence, hindered by limitations of existing benchmarks in scope and scale. To address this, we prese…