4 papers
VAM: Verbalized Action Masking for Controllable Exploration in RL Post-Training -- A Chess Case Study
Zhicheng Zhang, Ziyan Wang, Yali Du +1
Exploration remains a key bottleneck for reinforcement learning (RL) post-training of large language models (LLMs), where sparse feedback and large action spaces can lead to premat…
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models
Zhicheng Zhang, Ziyan Wang, Yali Du +1
Developing effective instruction-following policies in reinforcement learning remains challenging due to the reliance on extensive human-labeled instruction datasets and the diffic…
M3HF: Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed Quality
Ziyan Wang, Zhicheng Zhang, Fei Fang +1
Designing effective reward functions in multi-agent reinforcement learning (MARL) is a significant challenge, often leading to suboptimal or misaligned behaviors in complex, coordi…
Safe Multi-agent Reinforcement Learning with Natural Language Constraints
Ziyan Wang, Meng Fang, Tristan Tomilin +2
The role of natural language constraints in Safe Multi-agent Reinforcement Learning (MARL) is crucial, yet often overlooked. While Safe MARL has vast potential, especially in field…