Showing 2026Show all
3 papers · 1 filter
cs.LG2026
Train at Moving Edge: Online-Verified Prompt Selection for Efficient RL Training of Large Reasoning Model
Jiahao Wu, Ning Lu, Shengcai Liu +6
Reinforcement learning (RL) has become essential for post-training large language models (LLMs) in reasoning tasks. While scaling rollouts can stabilize training and enhance perfor…
cs.LG2026
DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation
Yang Zhou, Can Jin, Zihan Dong +7
Reinforcement learning improves the reasoning ability of large language models but remains costly and sample-inefficient, as many rollouts provide weak learning signals. Difficulty…
cs.RO2026
Seeing Farther and Smarter: Value-Guided Multi-Path Reflection for VLM Policy Optimization
Yanting Yang, Shenyuan Gao, Qingwen Bu +2
Solving complex, long-horizon robotic manipulation tasks requires a deep understanding of physical interactions, reasoning about their long-term consequences, and precise high-leve…