4 papers
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games
Tristan Maidment, JB Lanier, Chase McDonald +5
Recent work has established that regularized policy gradient methods such as PPO, when used in self-play, can match or exceed specialized game-theoretic algorithms for solving two-…
What metrics of participation balance predict outcomes of collaborative learning with a robot?
Yuya Asano, Diane Litman, Quentin King-Shepard +6
One of the keys to the success of collaborative learning is balanced participation by all learners, but this does not always happen naturally. Pedagogical robots have the potential…
Modeling Non-Cooperative Dialogue: Theoretical and Empirical Insights
Anthony Sicilia, Tristan Maidment, Pat Healy +1
Investigating cooperativity of interlocutors is central in studying pragmatics of dialogue. Models of conversation that only assume cooperative agents fail to explain the dynamics…
Domain-robust VQA with diverse datasets and methods but no target labels
Mingda Zhang, Tristan Maidment, Ahmad Diab +2
The observation that computer vision methods overfit to dataset specifics has inspired diverse attempts to make object recognition models robust to domain shifts. However, similar…