7 citations · 11 across the 2 of their papers we have counts for
4 papers
OptiDICE: Offline Policy Optimization via Stationary Distribution Correction Estimation
Jongmin Lee, Wonseok Jeon, Byung-Jun Lee +2
We consider the offline reinforcement learning (RL) setting where the agent aims to optimize the policy solely from the data without further environment interactions. In offline RL…
Regularized Inverse Reinforcement Learning
Wonseok Jeon, Chen-Yang Su, Paul Barde +3
Inverse Reinforcement Learning (IRL) aims to facilitate a learner's ability to imitate expert behavior by acquiring reward functions that explain the expert's decisions. Regularize…
Adversarial Soft Advantage Fitting: Imitation Learning without Policy Optimization
Paul Barde, Julien Roy, Wonseok Jeon +3
Adversarial Imitation Learning alternates between learning a discriminator -- which tells apart expert's demonstrations from generated ones -- and a generator's policy to produce t…
Scalable Multi-Agent Inverse Reinforcement Learning via Actor-Attention-Critic
Wonseok Jeon, Paul Barde, Derek Nowrouzezahrai +1
Multi-agent adversarial inverse reinforcement learning (MA-AIRL) is a recent approach that applies single-agent AIRL to multi-agent problems where we seek to recover both policies…