2 papers
cs.LG2026
RuleSmith: Multi-Agent LLMs for Automated Game Balancing
Ziyao Zeng, Chen Liu, Tianyu Liu +5
Game balancing is a longstanding challenge requiring repeated playtesting, expert intuition, and extensive manual tuning. We introduce RuleSmith, the first framework that achieves…
cs.LG2026
Your Group-Relative Advantage Is Biased
Fengkai Yang, Zherui Chen, Xiaohan Wang +10
Reinforcement Learning from Verifier Rewards (RLVR) has emerged as a widely used approach for post-training large language models on reasoning tasks, with group-based methods such…