2 papers
cs.LG2025
CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention
Qingbin Li, Rongkun Xue, Jie Wang +8
Recent advances in Reinforcement Learning with Verified Reward (RLVR) have driven the emergence of more sophisticated cognitive behaviors in large language models (LLMs), thereby e…
cs.MA2025
A Generalist Hanabi Agent
Arjun V Sudhakar, Hadi Nekoei, Mathieu Reymond +3
Traditional multi-agent reinforcement learning (MARL) systems can develop cooperative strategies through repeated interactions. However, these systems are unable to perform well on…