Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Learn More with Less: Uncertainty Consistency Guided Query Selection for RLVR
Hao Yi, Yulan Hu, Xin Li +3
Large Language Models (LLMs) have recently improved mathematical reasoning through Reinforcement Learning with Verifiable Reward (RLVR). However, existing RLVR algorithms require l…
cs.AI2025
Gradient Coupling: The Hidden Barrier to Generalization in Agentic Reinforcement Learning
Jingyu Liu, Xiaopeng Wu, Jingquan Peng +4
Reinforcement learning (RL) is a dominant paradigm for training autonomous agents, yet these agents often exhibit poor generalization, failing to adapt to scenarios not seen during…