Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Model-Based Exploration in Monitored Markov Decision Processes
Alireza Kazemipour, Simone Parisi, Matthew E. Taylor +1
A tenet of reinforcement learning is that the agent always observes rewards. However, this is not true in many realistic settings, e.g., a human observer may not always be availabl…
cs.LG2024
Beyond Optimism: Exploration With Partially Observable Rewards
Simone Parisi, Alireza Kazemipour, Michael Bowling
Exploration in reinforcement learning (RL) remains an open challenge. RL algorithms rely on observing rewards to train the agent, and if informative rewards are sparse the agent le…