17 papers
Multi-agent imitation learning with function approximation: Linear Markov games and beyond
Luca Viano, Till Freihaut, Emanuele Nevali +3
In this work, we present the first theoretical analysis of multi-agent imitation learning (MAIL) in linear Markov games where both the transition dynamics and each agent's reward f…
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference
Raphaël Baur, Yannick Metz, Maria Gkoulta +3
Reward learning typically relies on a single feedback type or combines multiple feedback types using manually weighted loss terms. Currently, it remains unclear how to jointly lear…
Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum
Navdeep Kumar, Tehila Dahan, Lior Cohen +4
We establish an optimal sample complexity of for obtaining an -optimal global policy using a single-timescale actor-critic (AC) algorithm in infinite-horizon disco…
Learning Acrobatic Flight from Preferences
Colin Merk, Ismail Geles, Jiaxu Xing +3
Preference-based reinforcement learning (PbRL) enables agents to learn control policies without requiring manually designed reward functions, making it well-suited for tasks where…
Aligning Language Models from User Interactions
Thomas Kleine Buening, Jonas Hübotter, Barna Pásztor +3
Multi-turn user interactions are among the most abundant data produced by language models, yet we lack effective methods to learn from them. While typically discarded, these intera…
On Feasible Rewards in Multi-Agent Inverse Reinforcement Learning
Till Freihaut, Giorgia Ramponi
Multi-agent Inverse Reinforcement Learning (MAIRL) aims to recover agent reward functions from expert demonstrations. We characterize the feasible reward set in Markov games, ident…