15 citations · 16 across the 9 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
VAM: Verbalized Action Masking for Controllable Exploration in RL Post-Training -- A Chess Case Study
Zhicheng Zhang, Ziyan Wang, Yali Du +1
Exploration remains a key bottleneck for reinforcement learning (RL) post-training of large language models (LLMs), where sparse feedback and large action spaces can lead to premat…
cs.LG2025
SocialJax: An Evaluation Suite for Multi-agent Reinforcement Learning in Sequential Social Dilemmas
Zihao Guo, Shuqing Shi, Richard Willis +3
Sequential social dilemmas pose a significant challenge in the field of multi-agent reinforcement learning (MARL), requiring environments that accurately reflect the tension betwee…