3 citations · 3 across the 6 of their papers we have counts for
6 papers
The Importance of Online Data: Understanding Preference Fine-tuning via Coverage
Yuda Song, Gokul Swamy, Aarti Singh +2
Learning from human preference data has emerged as the dominant paradigm for fine-tuning large language models (LLMs). The two most common families of techniques -- online reinforc…
Multi-Agent Imitation Learning: Value is Easy, Regret is Hard
Jingwu Tang, Gokul Swamy, Fei Fang +1
We study a multi-agent imitation learning (MAIL) problem where we take the perspective of a learner attempting to coordinate a group of agents based on demonstrations of an expert…
EvIL: Evolution Strategies for Generalisable Imitation Learning
Silvia Sapora, Gokul Swamy, Chris Lu +2
Often times in imitation learning (IL), the environment we collect expert demonstrations in and the environment we want to deploy our learned policy in aren't exactly the same (e.g…
The Virtues of Pessimism in Inverse Reinforcement Learning
David Wu, Gokul Swamy, J. Andrew Bagnell +2
Inverse Reinforcement Learning (IRL) is a powerful framework for learning complex behaviors from expert demonstrations. However, it traditionally requires repeatedly solving a comp…
Learning Shared Safety Constraints from Multi-task Demonstrations
Konwoo Kim, Gokul Swamy, Zuxin Liu +3
Regardless of the particular task we want them to perform in an environment, there are often shared safety constraints we want our agents to respect. For example, regardless of whe…
Game-Theoretic Algorithms for Conditional Moment Matching
Gokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell +1
A variety of problems in econometrics and machine learning, including instrumental variable regression and Bellman residual minimization, can be formulated as satisfying a set of c…