1 citations · 1 across the 5 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025★ 1 cited
CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention
Qingbin Li, Rongkun Xue, Jie Wang +8
Recent advances in Reinforcement Learning with Verified Reward (RLVR) have driven the emergence of more sophisticated cognitive behaviors in large language models (LLMs), thereby e…
cs.LG2025
Efficient Skill Discovery via Regret-Aware Optimization
He Zhang, Ming Zhou, Shaopeng Zhai +2
Unsupervised skill discovery aims to learn diverse and distinguishable behaviors in open-ended reinforcement learning. For existing methods, they focus on improving diversity throu…