13 papers
When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning
Luca Viano, Antoine Moulin, Audrey Huang +3
Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications from robotics to language model training. Standard approaches…
Multi-agent imitation learning with function approximation: Linear Markov games and beyond
Luca Viano, Till Freihaut, Emanuele Nevali +3
In this work, we present the first theoretical analysis of multi-agent imitation learning (MAIL) in linear Markov games where both the transition dynamics and each agent's reward f…
Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation
Ziyad Sheebaelhamd, Luca Viano, Volkan Cevher +1
This work investigates multi-objective imitation learning: the problem of recovering policies that lie on the Pareto front given demonstrations from multiple Pareto-optimal experts…
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution
Adam Barla, Emanuele Nevali, Luca Viano +1
We introduce PEPO (Pessimistic Ensemble based Preference Optimization), a single-step Direct Preference Optimization (DPO)-like algorithm to mitigate the well-known over-optimizati…
Optimistically Optimistic Exploration for Provably Efficient Infinite-Horizon Reinforcement and Imitation Learning
Antoine Moulin, Gergely Neu, Luca Viano
We study the problem of reinforcement learning in infinite-horizon discounted linear Markov decision processes (MDPs), and propose the first computationally efficient algorithm ach…
Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback
Amirhossein Afsharrad, Ruida Zhou, Luca Viano +2
Reward modeling is crucial for aligning large language models with human preferences, yet current approaches lack a principled mathematical framework for leveraging ordinal prefere…