2 papers
cs.AI2026
Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning
Heyang Jiang, Henry Liu, Baharan Mirzasoleiman
Reinforcement learning with verifiable rewards (RLVR) has emerged as a highly effective framework for improving LLM reasoning, with methods such as GRPO among its most successful i…
cs.CL2025
IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation
Johannes Schmitt, Gergely Bérczi, Jasper Dekoninck +57
As the mathematical capabilities of large language models (LLMs) improve, it becomes increasingly important to evaluate their performance on research-level tasks at the frontier of…