22 citations · 97 across the 38 of their papers we have counts for
7 papers · 1 filter
SimuDICE: Offline Policy Optimization Through World Model Updates and DICE Estimation
Catalin E. Brita, Stephan Bongers, Frans A. Oliehoek
In offline reinforcement learning, deriving an effective policy from a pre-collected set of experiences is challenging due to the distribution mismatch between the target policy an…
Navigating Trade-offs: Policy Summarization for Multi-Objective Reinforcement Learning
Zuzanna Osika, Jazmin Zatarain-Salazar, Frans A. Oliehoek +1
Multi-objective reinforcement learning (MORL) is used to solve problems involving multiple objectives. An MORL agent must make decisions based on the diverse signals provided by di…
Communicating with Speakers and Listeners of Different Pragmatic Levels
Kata Naszadi, Frans A. Oliehoek, Christof Monz
This paper explores the impact of variable pragmatic competence on communicative success through simulating language learning and conversing between speakers and listeners with dif…
Online Planning in POMDPs with State-Requests
Raphael Avalos, Eugenio Bargiacchi, Ann Nowé +2
In key real-world problems, full state information is sometimes available but only at a high cost, like activating precise yet energy-intensive sensors or consulting humans, thereb…
Inverse Concave-Utility Reinforcement Learning is Inverse Game Theory
Mustafa Mert Çelikok, Frans A. Oliehoek, Jan-Willem van de Meent
We consider inverse reinforcement learning problems with concave utilities. Concave Utility Reinforcement Learning (CURL) is a generalisation of the standard RL objective, which em…
Policy Space Response Oracles: A Survey
Ariyan Bighashdel, Yongzhao Wang, Stephen McAleer +2
Game theory provides a mathematical way to study the interaction between multiple decision makers. However, classical game-theoretic analysis is limited in scalability due to the l…