116 citations · 176 across the 13 of their papers we have counts for
13 papers
The Importance of Online Data: Understanding Preference Fine-tuning via Coverage
Yuda Song, Gokul Swamy, Aarti Singh +2
Learning from human preference data has emerged as the dominant paradigm for fine-tuning large language models (LLMs). The two most common families of techniques -- online reinforc…
Hybrid Reinforcement Learning from Offline Observation Alone
Yuda Song, J. Andrew Bagnell, Aarti Singh
We consider the hybrid reinforcement learning setting where the agent has access to both offline data and online interactive access. While Reinforcement Learning (RL) research typi…
The Virtues of Pessimism in Inverse Reinforcement Learning
David Wu, Gokul Swamy, J. Andrew Bagnell +2
Inverse Reinforcement Learning (IRL) is a powerful framework for learning complex behaviors from expert demonstrations. However, it traditionally requires repeatedly solving a comp…
The Virtues of Laziness in Model-based RL: A Unified Objective and Algorithms
Anirudh Vemula, Yuda Song, Aarti Singh +2
We propose a novel approach to addressing two fundamental challenges in Model-based Reinforcement Learning (MBRL): the computational expense of repeatedly finding a good policy in…
Game-Theoretic Algorithms for Conditional Moment Matching
Gokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell +1
A variety of problems in econometrics and machine learning, including instrumental variable regression and Bellman residual minimization, can be formulated as satisfying a set of c…
On the Effectiveness of Iterative Learning Control
Anirudh Vemula, Wen Sun, Maxim Likhachev +1
Iterative learning control (ILC) is a powerful technique for high performance tracking in the presence of modeling errors for optimal control applications. There is extensive prior…