3 papers
cs.LG2025
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models
Zhicheng Zhang, Ziyan Wang, Yali Du +1
Developing effective instruction-following policies in reinforcement learning remains challenging due to the reliance on extensive human-labeled instruction datasets and the diffic…
cs.MA2025
M3HF: Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed Quality
Ziyan Wang, Zhicheng Zhang, Fei Fang +1
Designing effective reward functions in multi-agent reinforcement learning (MARL) is a significant challenge, often leading to suboptimal or misaligned behaviors in complex, coordi…
cs.LG2024
Natural Language Reinforcement Learning
Xidong Feng, Bo Liu, Yan Song +7
Artificial intelligence progresses towards the "Era of Experience," where agents are expected to learn from continuous, grounded interaction. We argue that traditional Reinforcemen…