Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Understanding R1-Zero-Like Training: A Critical Perspective
Zichen Liu, Changyu Chen, Wenjun Li +5
DeepSeek-R1-Zero has shown that reinforcement learning (RL) at scale can directly enhance the reasoning capabilities of LLMs without supervised fine-tuning. In this work, we critic…
cs.LG2025
Improving Environment Novelty Quantification for Effective Unsupervised Environment Design
Jayden Teoh, Wenjun Li, Pradeep Varakantham
Unsupervised Environment Design (UED) formalizes the problem of autocurricula through interactive training between a teacher agent and a student agent. The teacher generates new tr…