Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games
Tristan Maidment, JB Lanier, Chase McDonald +5
Recent work has established that regularized policy gradient methods such as PPO, when used in self-play, can match or exceed specialized game-theoretic algorithms for solving two-…
cs.LG2024
Skill Set Optimization: Reinforcing Language Model Behavior via Transferable Skills
Kolby Nottingham, Bodhisattwa Prasad Majumder, Bhavana Dalvi Mishra +3
Large language models (LLMs) have recently been used for sequential decision making in interactive environments. However, leveraging environment reward signals for continual LLM ac…