1 citations · 1 across the 12 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning
Shujin Wu, Cheng Qian, Xiusi Chen +1
Test-time scaling through iterative self-evolution with environment feedback, as demonstrated by AlphaEvolve, shows remarkable performance gains. We hypothesize that the success of…
cs.LG2026
Decoding the Critique Mechanism in Large Reasoning Models
Hoang Phan, Quang H. Nguyen, Hung T. Q. Le +3
Large Reasoning Models (LRMs) exhibit backtracking and self-verification mechanisms that enable them to revise intermediate steps and reach correct solutions, yielding strong perfo…
cs.LG2026
How Far Can Unsupervised RLVR Scale LLM Training?
Bingxiang He, Yuxin Zuo, Zeyuan Liu +18
Unsupervised reinforcement learning with verifiable rewards (URLVR) offers a pathway to scale LLM training beyond the supervision bottleneck by deriving rewards without ground trut…