2 papers
cs.LG2026
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games
Tristan Maidment, JB Lanier, Chase McDonald +5
Recent work has established that regularized policy gradient methods such as PPO, when used in self-play, can match or exceed specialized game-theoretic algorithms for solving two-…
cs.HC2024
What metrics of participation balance predict outcomes of collaborative learning with a robot?
Yuya Asano, Diane Litman, Quentin King-Shepard +6
One of the keys to the success of collaborative learning is balanced participation by all learners, but this does not always happen naturally. Pedagogical robots have the potential…