5 papers
Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
Mickel Liu, Liwei Jiang, Yancheng Liang +4
Conventional large language model (LLM) safety alignment relies on a reactive, disjoint loop: attackers exploit a static model, then defenders patch exposed vulnerabilities. This s…
Learn to Match: Two-Sided Matching with Temporally Extended Feedback
Haijing Zong, Yancheng Liang, Boyang Zhou +1
Two-sided matching markets often involve information that unfolds over time through interviews, repeated interaction, learning, and separation. Existing matching models typically r…
Improving Human-AI Coordination through Online Adversarial Training and Generative Models
Paresh Chaudhary, Yancheng Liang, Daphne Chen +2
Being able to cooperate with diverse humans is an important component of many economically valuable AI tasks, from household robotics to autonomous driving. However, generalizing t…
Cross-environment Cooperation Enables Zero-shot Multi-agent Coordination
Kunal Jha, Wilka Carvalho, Yancheng Liang +3
Zero-shot coordination (ZSC), the ability to adapt to a new partner in a cooperative task, is a critical component of human-compatible AI. While prior work has focused on training…
Learning to Cooperate with Humans using Generative Agents
Yancheng Liang, Daphne Chen, Abhishek Gupta +2
Training agents that can coordinate zero-shot with humans is a key mission in multi-agent reinforcement learning (MARL). Current algorithms focus on training simulated human partne…