chain-of-thought reasoning 1emergent behavior 1large language models 1scaling laws 1zero-shot reinforcement learning 1
From the 1 of 13 linked papers with an AI index.
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
When Sharpening Becomes Collapse: Sampling Bias and Semantic Coupling in RL with Verifiable Rewards
Mingyuan Fan, Weiguang Han, Daixin Wang +3
Reinforcement Learning with Verifiable Rewards (RLVR) is a central paradigm for turning large language models (LLMs) into reliable problem solvers, especially in logic-heavy domain…
cs.LG2025
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward
Xinyu Tang, Zhenduo Zhang, Yurou Liu +4
Recent advances in large reasoning models have leveraged reinforcement learning with verifiable rewards (RLVR) to improve reasoning capabilities. However, scaling these methods typ…
cs.LG2025
IceBerg: Debiased Self-Training for Class-Imbalanced Node Classification
Zhixun Li, Dingshuo Chen, Tong Zhao +5
Graph Neural Networks (GNNs) have achieved great success in dealing with non-Euclidean graph-structured data and have been widely deployed in many real-world applications. However,…